Papers with language modeling

281 papers
Dive into Deep Learning for Natural Language Processing (D19-2)

Copied to clipboard

Challenge: GluonNLP is a powerful new toolkit that automates the most laborious aspects of deep learning for NLP.
Approach: This hands-on tutorial demonstrates how to scale unsupervised pre-training techniques with Apache MXNet and GluonNLP.
Outcome: This hands-on tutorial examines the challenges of scaling these models and algorithms effectively with Apache MXNet and GluonNLP.
Latent Structure Models for Natural Language Processing (P19-4)

Copied to clipboard

Challenge: Latent structure models are a powerful tool for compositional data modeling and pipelines.
Approach: This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations .
Outcome: This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations .
What BERT Is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models (2020.tacl-1)

Copied to clipboard

Challenge: Pretraining by language modeling has become popular but we have yet to understand what language models learn during that process.
Approach: They propose diagnostics that ask questions about information used by language models for generating predictions in context.
Outcome: The proposed diagnostics can be used to study the popular BERT model . they show that the model can distinguish good from bad completions, but struggles with inference and role-based event prediction.
Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token Embeddings (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have determined that the learned token embeddings of large-scale neural language models are degenerated to be anisotropic with a narrow-cone shape.
Approach: They propose a method to degenerate the learning gradient for rare token embeddings by gating the specific part of the gradient for all tokens during training stage.
Outcome: The proposed method improves the performance of the models but lacks the training dynamics needed to solve the representation degeneration problem.
Building Hierarchically Disentangled Language Models for Text Generation with Named Entities (2020.coling-main)

Copied to clipboard

Challenge: Named entities pose a unique challenge to traditional methods of language modeling.
Approach: They propose a Hierarchically Disentangled Model for named entities in cooking recipes using a dataset from several publicly available online sources.
Outcome: The proposed model is based on 158,473 cooking recipes from public sources.
Deep RNNs Encode Soft Hierarchical Syntax (P18-2)

Copied to clipboard

Challenge: Existing studies show that syntactic information is useful for a wide variety of NLP tasks.
Approach: They propose to use word-level representations to learn internal representations that capture soft hierarchical notions of syntax from highly varied supervision.
Outcome: The proposed model encodes significant amounts of syntax even without explicit supervision.
Efficient Content-Based Sparse Attention with Routing Transformers (2021.tacl-1)

Copied to clipboard

Challenge: Self-attention suffers from quadratic computation and memory requirements with respect to sequence length . despite its effectiveness, self-attention models suffer from quadratic computation and a limited set of locations .
Approach: They propose to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest.
Outcome: The proposed model outperforms similar sparse attention models on language modeling and image generation on Wikitext-103 .
Towards Codec-LM Co-design for Neural Codec Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Neural codec language models (or codec LMs) are emerging as a powerful framework for text-to-speech (TTS) despite the close interdependence of codecs and LM, research on codec and lms has largely remained siloed.
Approach: They propose a frame-wise codec encoder that improves both LM log-likelihood and TTS metrics . they also propose LM codebook level dropout to efficiently navigate a portion of codec-LM design space .
Outcome: The proposed codec-LM co-design improves intelligibility, audio quality and speaker control compared to a siloed baseline.
Scaling Parameter-Constrained Language Models with Quality Data (2024.emnlp-industry)

Copied to clipboard

Challenge: Scaling laws in language modeling quantify training loss as a function of dataset size and model parameters, but neglect the critical role of data quality in model generalization.
Approach: They propose to use effective training tokens as a combination of text diversity and syntheticity as measured by a teacher model to calculate scaling laws.
Outcome: The proposed term effective training tokens is a combination of two readily-computed indicators of text diversity and syntheticity as measured by a teacher model.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling (N19-4)

Copied to clipboard

Challenge: OpenNMT is a community-built toolkit written in multiple languages with an emphasis on extensibility.
Approach: They propose to use PyTorch to train custom sequence models for translation, summarization, language modeling, and other tasks.
Outcome: The proposed toolkit is fast, extensible, and useful for both research and production.
The Geometry of Multilingual Language Model Representations (2022.emnlp-main)

Copied to clipboard

Challenge: XLM-R models encode language-sensitive information in each language, allowing them to extract features for downstream tasks and cross-lingual transfer learning.
Approach: They evaluate how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language.
Outcome: The proposed model can extract features for downstream tasks and cross-lingual transfer learning.
Returning to the Start: Generating Narratives with Related Endpoints (2024.naacl-short)

Copied to clipboard

Challenge: RENarGen generates closed narratives by ensuring the first and last sentences are related and then infilling the middle sentences.
Approach: They propose a novel novel novel that generates closed narratives by ensuring the first and last sentences are related and then infilling the middle sentences.
Outcome: The proposed paradigm generates closed narratives by ensuring the first and last sentences are related and then infilling the middle sentences.
Controlled Text Generation Using Dictionary Prior in Variational Autoencoders (2022.findings-acl)

Copied to clipboard

Challenge: Variational autoencoders (VAEs) have been widely applied in text generation tasks, but they suffer from insufficient representation capacity and poor controllability.
Approach: They propose a data-driven prior that has expressivity and controllability.
Outcome: The proposed prior enjoys expressivity and controllability and can be used in language modeling and controlled text generation.
Attention Alignment and Flexible Positional Embeddings Improve Transformer Length Extrapolation (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for length extrapolation are tailored for natural language modeling, a task known to have strong recency bias.
Approach: They propose two attention alignment strategies to improve T5's long-context utilization capability without fine-tuning.
Outcome: The proposed methods improve the long-context utilization capability of T5 on language modeling, retrieval, multi-document question answering, and code completion tasks without any fine-tuning.
Concealed Data Poisoning Attacks on NLP Models (2021.naacl-main)

Copied to clipboard

Challenge: In contrast, adversarial attacks can cause model errors by modifying inputs, such as the universal triggers attack.
Approach: They propose a data poisoning attack that allows an adversary to control model predictions whenever a desired trigger phrase is present in the input.
Outcome: The proposed attack can cause model errors by modifying inputs, but it can also cause extra human annotation.
Interpreting Language Models with Contrastive Explanations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing explanation methods conflate evidence for various features to predict a token . existing explanation methods are less interpretable for human understanding .
Approach: They propose to explain language models contrastively by looking for salient input tokens that explain why the model predicted one token instead of another.
Outcome: The proposed explanations are better than non-contrastive explanations for language models . they show that contrastive explanations improve simulability for human observers .
Morphology Matters: A Multilingual Language Modeling Analysis (2021.tacl-1)

Copied to clipboard

Challenge: Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model.
Approach: They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model.
Outcome: The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling.
Generalization in Generation: A closer look at Exposure Bias (D19-56)

Copied to clipboard

Challenge: Autoregressive generative models are often criticized for using ground-truth contexts at training time but generated ones at test time.
Approach: They propose that generalization is the underlying property to address and propose unconditional generation as its fundamental benchmark.
Outcome: The proposed model is generalized and can handle true and generated contexts.
Few-Shot NLG with Pre-Trained Language Model (2020.acl-main)

Copied to clipboard

Challenge: Neural-based approaches to natural language generation are data-hungry and difficult to adopt in real-world applications.
Approach: They propose a task of few-shot natural language generation from structured data or knowledge to generate coherent sentences from input data and language modeling to compose coherent sentences.
Outcome: The proposed approach outperforms the strongest baseline approach by over 8.0 BLEU points improvement.
Code Representation Pre-training with Complements from Program Executions (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing languages have syntactic representations of code to improve code intelligence, but they are difficult to learn from code.
Approach: They propose to embed dynamic information of programs revealed by their test cases into feature representations of code as complements.
Outcome: The proposed method yields 6%/19% mAP improvements over its masked language modeling counterparts.
Cyclical Annealing Schedule: A Simple Approach to Mitigating KL Vanishing (N19-1)

Copied to clipboard

Challenge: Variational autoencoders (VAEs) with an auto-regressive decoder have been applied for many natural language processing tasks.
Approach: They propose a cyclical annealing schedule which repeats the process of increasing multiple times to learn more meaningful latent codes progressively by leveraging previous learning cycles as warm re-restart.
Outcome: The proposed method improves on a broad range of NLP tasks, including language modeling, dialog response generation and semi-supervised text classification.
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages (2025.acl-srw)

Copied to clipboard

Challenge: Low-resource languages (LRLs) face significant challenges in natural language processing due to limited data.
Approach: They evaluate adapter-based methods for adapting mLMs to low-resource languages . they use unstructured text and structured knowledge from ConceptNet to evaluate adapters .
Outcome: The proposed methods outperform large language models and LLaMA-3 and deepSeek-R1 models on low training data.
Kronecker Decomposition for GPT Compression (2022.acl-short)

Copied to clipboard

Challenge: GPT is an auto-regressive Transformer-based pre-trained language model . but its huge size can be prohibitive for deploying on low capacity devices .
Approach: They use a Kronecker decomposition technique to compress GPT models . they use ILKD to refine the model on downstream tasks .
Outcome: The proposed model outperforms the existing DistilGPT2 model on language modeling and general language understanding evaluation benchmark tasks.
Linearizing Transformer with Key-Value Memory (2022.emnlp-main)

Copied to clipboard

Challenge: Efficient transformer variants with linear time complexity have been developed to mitigate the quadratic computational overhead of the vanilla transformer.
Approach: They propose a linear time complexity transformer variant that reduces the quadratic computational overhead of the vanilla transformer by using a recurrent-style incremental computation similar to kernel-based transformers.
Outcome: The proposed method reduces the performance gap while achieving the same efficiency even with short generation.
Character-Based Neural Networks for Sentence Pair Modeling (N18-2)

Copied to clipboard

Challenge: Sentence pair modeling is critical for many NLP tasks, such as paraphrase identification and semantic textual similarity.
Approach: They propose to use subwords to represent sentences without pretrained word embeddings . they find that subword models can achieve new state-of-the-art results without pretraining .
Outcome: The proposed models can achieve state-of-the-art results on two social media datasets and competitive results on news data for paraphrase identification.
Riemannian Normalizing Flow on Variational Wasserstein Autoencoder for Text Modeling (N19-1)

Copied to clipboard

Challenge: Empirical experiments show that our model learns latent distributions that respect latent space geometry and is able to generate sentences that are more diverse.
Approach: They propose a Variational Wasserstein Autoencoder with Riemannian Normalizing Flow to solve this problem by transforming a latent variable into a space that respects the geometric characteristics of input space.
Outcome: Empirical results show that the proposed model avoids KLvanishing and has better performance in language modeling, likelihood approximation, and text generation tasks.
More room for language: Investigating the effect of retrieval on language models (2024.naacl-short)

Copied to clipboard

Challenge: Retrieval-augmented language models are a promising alternative to standard pretraining, but little attention has been put into understanding what this type of training scheme does to the underlying language model when analyzed as a standalone -separated from the overall retrieval pipeline.
Approach: They propose an ‘ideal retrieval’ methodology to study these models in a fully controllable setting and propose a retrieval augmentation methodology to examine their effects.
Outcome: The proposed model saves substantially less world knowledge in their weights, but is worse at comprehending global context.
On-Device Neural Language Model Based Word Prediction (C18-2)

Copied to clipboard

Challenge: Currently, on-device keyboards have limited memory and response time for word prediction . a proposed on-device neural language model based word prediction method is available for mobile devices .
Approach: They propose an on-device neural language model based word prediction method that optimizes run-time memory and provides a real-time prediction environment.
Outcome: The proposed model outperforms existing methods for word prediction in keystroke savings and word prediction rate and has been commercialized.
On the Relation between Linguistic Typology and (Limitations of) Multilingual Language Modeling (D18-1)

Copied to clipboard

Challenge: a key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language.
Approach: They propose to use a full-vocabulary setup to test the performance of language modeling (LM) on 50 typologically diverse languages.
Outcome: The proposed language modeling task is based on a full vocabulary setup focused on word-level prediction on 50 typologically diverse languages.
Language Models Get a Gender Makeover: Mitigating Gender Bias with Few-Shot Data Interventions (2023.acl-short)

Copied to clipboard

Challenge: Existing approaches to de-bias pre-trained large language models focus on changes to training regime, but this is not feasible.
Approach: They propose to de-bias a pre-trained model by fine-tuning it on only 10 examples . they show that the technique performs better than competitive baselines .
Outcome: The proposed method performs better than competitive state-of-the-art baselines with minimal loss in language modeling ability.
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? A Petroglyph Revisited (2025.findings-acl)

Copied to clipboard

Challenge: Autoregressive Transformer language models do not require explicit positional encodings (PEs) this is because a cascade of (permutation invariant) set processors collectively exhibit sequence-sensitive behavior in the autoregressively setting.
Approach: They propose to explain why autoregressive Transformers require explicit positional encodings (PEs) this property has been known since early efforts adopting the Transformer for language modeling .
Outcome: The proposed model can distinguish sequences with permuted tokens without the need for explicit PEs.
FlashBack: Efficient Retrieval-Augmented Language Modeling for Fast Inference (2025.findings-acl)

Copied to clipboard

Challenge: Retrieval-Augmented Language Modeling (RALM) is a popular approach for large language models.
Approach: They propose a modular RALM that integrates large language models with documents from an external corpus to improve inference efficiency.
Outcome: The proposed method improves inference efficiency with appending context pattern while maintaining decent performance after fine-tuning by Low-Rank Adaption.
Finding a Needle in the Adversarial Haystack: A Targeted Paraphrasing Approach For Uncovering Edge Cases with Minimal Distribution Distortion (2024.eacl-long)

Copied to clipboard

Challenge: Adversarial attacks against Language models (LMs) are a significant concern.
Approach: They propose an approach to automatically learn a policy to generate challenging examples that improve the model’s performance.
Outcome: The proposed approach outperforms baselines and exhibits generalizability across classifiers and datasets.
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference (2025.naacl-short)

Copied to clipboard

Challenge: Recent studies have shown that sharing key-value (KV) cache across layers is effective in efficient inference of large language models.
Approach: They propose a unified framework that covers several recent methods and their novel variants to investigate cross-layer KV sharing.
Outcome: The proposed framework achieves higher throughput and better performance when reducing the size of the key-value cache by 2 while maintaining competitive performance.
Visualizing the Relationship Between Encoded Linguistic Information and Task Performance (2022.findings-acl)

Copied to clipboard

Challenge: Recent studies show that encoding more syntactic information does not lead to better performance.
Approach: They propose a method to optimize pareto-optimal models by formalizing it as a multi-objective optimization problem.
Outcome: The proposed method is better than a baseline method on two NLP tasks.
APo-VAE: Text Generation in Hyperbolic Space (2021.naacl-main)

Copied to clipboard

Challenge: Existing models that learn embeddings only in Euclidean vector space do not account for such structural property of language.
Approach: They propose a Poincare Variational Autoencoder to capture latent hierarchies in hyperbolic space . they propose enabling adversarial learning procedures to empower robust model training .
Outcome: The proposed model outperforms existing models in a hyperbolic latent space . it captures latent language hierarchies in hyperbolical space and is robust to training .
Internal and External Impacts of Natural Language Processing Papers (2025.acl-short)

Copied to clipboard

Challenge: a new study examines the impact of NLP research published in top-tier conferences from 1979 to 2024 . language modeling has the widest internal and external influence, while linguistic foundations have lower impacts .
Approach: They analyze citations from research articles and external sources to determine how NLP topics are consumed internally and externally.
Outcome: The findings show that language modeling has the widest internal and external influence . ethics, bias, and fairness show significant attention in policy documents with fewer academic citations .
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models (2024.acl-demos)

Copied to clipboard

Challenge: Efficient fine-tuning of large language models requires non-trivial efforts to implement these methods on different models.
Approach: They propose a framework that democratizes the fine-tuning of large language models by integrating a suite of efficient training methods into one framework.
Outcome: The proposed framework is able to scale to 100+ LLMs without coding and receives over 25,000 stars and 3,000 forks.
Director: Generator-Classifiers For Supervised Language Modeling (2022.aacl-main)

Copied to clipboard

Challenge: Current language models achieve low perplexity but their resulting generations still suffer from toxic responses, repetitiveness, and contradictions.
Approach: They propose a new language model architecture that uses a language modeling and a classification head for each output token.
Outcome: The proposed model outperforms existing model guiding approaches in terms of accuracy and efficiency.
Empirical Sufficiency Lower Bounds for Language Modeling with Locally-Bootstrapped Semantic Structures (2023.starsem-1)

Copied to clipboard

Challenge: a recent attempt at language modeling with predicted semantic structure failed to establish empirical lower bounds on what could have made the attempt successful.
Approach: They propose a concise binary vector representation of semantic structure at the lexical level and evaluate how good an incremental tagger needs to be to achieve better-than-baseline performance.
Outcome: The proposed model can achieve better-than-baseline performance without losing its main advantages and lower bounds on prediction quality can't be established via a single score alone.
Great Memory, Shallow Reasoning: Limits of kNN-LMs (2025.naacl-short)

Copied to clipboard

Challenge: Existing models trained on poor quality data have shown strong performance in language modeling and some downstream benchmarks.
Approach: They evaluate kNN-LMs on a diverse set of tasks and evaluate their performance.
Outcome: The proposed extension could improve on a variety of tasks, but it fails to perform on reasoning tasks that require integrating multiple pieces of information.
Cascaded Head-colliding Attention (2021.acl-long)

Copied to clipboard

Challenge: Existing frameworks for natural language processing ignore interactions among different heads, which wastes the capacity of the model.
Approach: They propose a model which explicitly models interactions between attention heads through a hierarchical variational distribution.
Outcome: The proposed model outperforms the baseline model on Wikitext-103 and WMT14 EN-DE on language modeling and translation tasks.
Revisiting Representation Degeneration Problem in Language Modeling (2020.findings-emnlp)

Copied to clipboard

Challenge: Language modeling is a fundamental task in natural language processing, applications include machine translation, image captioning and speech recognition.
Approach: They propose a cosine regularization method to solve the representation degeneration problem by analyzing the limitations of the proposed method and then propose an alternative regularization technique to tackle the problem.
Outcome: The proposed method is effective in language modeling and image captioning.
Neurocache: Efficient Vector Retrieval for Long-range Language Modeling (2024.naacl-long)

Copied to clipboard

Challenge: Recent research shows that retrieval-augmented models with shorter contexts (4K tokens) can match the performance of models with longer contexts (16K/32K token)
Approach: They introduce an approach to extend the effective context size of large language models by using an external vector cache to store past states.
Outcome: The proposed method improves on models trained from scratch and pre-trained models.
ORBIT: Cost-Effective Dataset Curation for Large Language Model Domain Adaptation with an Astronomy Case Study (2025.findings-acl)

Copied to clipboard

Challenge: General-purpose models lack depth for expert-level tasks because of limited domain-specific information.
Approach: They propose a method for curating domain-specific datasets from noisy web sources to improve model performance.
Outcome: The proposed model outperforms the baseline model on the astronomy benchmark and on the AstroBench.
Human Language Modeling (2022.findings-acl)

Copied to clipboard

Challenge: Existing language modeling models treat text sequences as if they were created independently.
Approach: They propose a hierarchical extension to the language modeling problem whereby a human-level exists to connect sequences of documents and capture the notion that human language is moderated by changing human states.
Outcome: The proposed model outperforms the current state-of-the-art in terms of language modeling and fine-tuning for 4 downstream tasks spanning document- and user-levels.
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training (2024.findings-naacl)

Copied to clipboard

Challenge: Recent advances in language modeling have led to the emergence of large language models capable ofvarious natural language processing tasks.
Approach: They propose a multi-instructional training approach that integrates a large language model with a speech encoder to harness the capabilities of LLMs for speech recognition and beyond.
Outcome: The proposed model can be trained and aligned with a multilingual LLM on 1900 hours of transcribed data from 139 languages.
Stateful Memory-Augmented Transformers for Efficient Dialogue Modeling (2024.findings-eacl)

Copied to clipboard

Challenge: Existing Transformers models are computationally expensive for long context inputs.
Approach: They propose a transformer that can interchange information between memory states and context . they evaluate the efficiency of their model on three dialogue datasets and two language datasets .
Outcome: The proposed model is compatible with existing transformer models and can preserve dialogue history information.
Learning to Represent Image and Text with Denotation Graph (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in learning representations of visual and language information have been a problem with many applications.
Approach: They propose to extract visual expressions from images aligned with linguistic expressions that describe the images to learn representations from implicit expressions.
Outcome: The proposed representations lead to stronger empirical results on downstream tasks of cross-modal image retrieval, referring expression, and compositional attribute-object recognition.
Pivotal Role of Language Modeling in Recommender Systems: Enriching Task-specific and Task-agnostic Representation Learning (2023.acl-long)

Copied to clipboard

Challenge: Recent studies have proposed unified user modeling frameworks that leverage user behavior data from various applications.
Approach: They propose to use user behavior sequences as plain text to represent rich information in any domain or system without losing generality.
Outcome: The proposed frameworks achieve excellent results on diverse recommendation tasks and can be used on unseen domains and services.
Self-Normalization Properties of Language Modeling (C18-1)

Copied to clipboard

Challenge: Existing methods to reduce run-times for language models with large word vocabularies are based on noise contrastive estimation (NCE)
Approach: They propose to use noise-constrained noise-based models to approximate the normalized probability of a class without having to compute the partition function.
Outcome: The proposed model outperforms softmax-based models in a variety of NLP tasks and is based on the noise-constrained noise-constant estimation properties.
TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented Dialogue (2020.emnlp-main)

Copied to clipboard

Challenge: Existing pre-trained language models with self-attention encoder architectures are less useful in practice.
Approach: They propose to use user and system tokens to model dialogue behavior during pre-training . they propose a contrastive objective function to simulate the response selection task .
Outcome: The proposed model outperforms baseline models on four downstream tasks . it also has a few-shot ability that can mitigate the data scarcity problem .
Mukayese: Turkish NLP Strikes Back (2022.findings-acl)

Copied to clipboard

Challenge: Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the researched class.
Approach: They propose a set of NLP benchmarks for the Turkish language that contains several NLP tasks.
Outcome: The proposed benchmarks outperform previous work significantly in the Turkish language.
DS-TOD: Efficient Domain Specialization for Task-Oriented Dialog (2022.findings-acl)

Copied to clipboard

Challenge: Recent work shows that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining.
Approach: They propose a resource-efficient and modular domain specialization by means of domain adapters in which domain knowledge is encoded.
Outcome: The proposed framework extracts domain-specific terms and then uses them to build DomainCC and DomainReddit resources based on masked language modeling and response selection objectives.
Efficient Sequence Learning with Group Recurrent Networks (N18-1)

Copied to clipboard

Challenge: Recurrent neural networks have achieved state-of-the-art results in many artificial intelligence tasks, such as language modeling, neural machine translation and speech recognition.
Approach: They propose an efficient architecture to improve the efficiency of such RNN model training by adopting the group strategy for recurrent layers while exploiting the representation rearrangement strategy between layers as well as time steps.
Outcome: The proposed architecture achieves comparable or better accuracy compared with baselines, with a much smaller number of parameters and at a lower computational cost.
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in image tokenizers have enabled text-to-image generation using auto-regressive methods, but these methods lack pre-trained language models for text-based models.
Approach: They adapt a pre-trained language model for auto-regressive text-to-image generation and show that pre-train language models offer limited help.
Outcome: The proposed model is compared with a pre-trained language model and shows that it is no more effective than random initialized models.
In-Context Retrieval-Augmented Language Models (2023.tacl-1)

Copied to clipboard

Challenge: Existing RALM methods focus on modifying the LM architecture to facilitate incorporation of external information, complicating deployment.
Approach: They propose to condition a language model on relevant documents from a grounding corpus during generation by conditioning on external knowledge sources.
Outcome: The proposed method significantly improves language modeling performance and provides natural source attribution mechanism.
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets (2021.acl-long)

Copied to clipboard

Challenge: Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measurements of harms.
Approach: They apply a measurement modeling lens to inventory pitfalls that threaten benchmarks' validity as measurement models for stereotyping.
Outcome: The proposed benchmarks lack clarity and assumptions that affect how they conceptualize and operationalize stereotyping.
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars.
Approach: They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources.
Outcome: The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios.
Probing via Prompting (2022.naacl-main)

Copied to clipboard

Challenge: Pre-trained language models have increased the performance of data-driven natural language processing (NLP) models on a wide variety of tasks.
Approach: They propose a model-free approach to probing via prompting which formulates probing as a prompting task and combine pruning to analyze where the model stores the linguistic information in its architecture.
Outcome: The proposed approach extracts information from pre-trained models while learning much less on its own.
Neural Syntactic Generative Models with Exact Marginalization (N18-1)

Copied to clipboard

Challenge: Recent models have added structure to recurrent neural networks at the cost of giving up exact inference, or using soft structure instead of latent variables.
Approach: They propose a syntactic generative model with exact marginalization that supports dependency parsing and language modeling.
Outcome: The proposed models achieve state-of-the-art for supervised dependency parsing and language modeling.
Online Back-Parsing for AMR-to-Text Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Abstract meaning representation (AMR) is a semantic graph representation that abstracts meaning away from a sentence.
Approach: They propose a decoder that back predicts projected AMR graphs on target sentences . their results show superiority over previous state-of-the-art decoded graph Transformer .
Outcome: The proposed model outperforms the state-of-the-art model on two AMR benchmarks.
Grounded Compositional Outputs for Adaptive Language Modeling (2020.emnlp-main)

Copied to clipboard

Challenge: Language models are a key component of natural language processing, but their size is a problem because they are typically trained with a closed output vocabulary derived from the training data.
Approach: They propose a fully compositional output embedding layer for language models that is grounded in semantically related words and free-text definitions.
Outcome: The proposed model outperforms state-of-the-art methods and adaptation approaches on cross-domain modeling and cross-learning tasks.
Incorporating Stylistic Lexical Preferences in Generative Language Models (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in language modeling have resulted in powerful generation models, but their style is implicitly dependent on the training data and cannot emulate a specific target style.
Approach: They propose an approach to induce certain target-author attributes by incorporating continuous multi-dimensional lexical preferences of an author into generative language models.
Outcome: The proposed model generates text that aligns with a given target author’s lexical style and is competitive with baselines.
Tree Transformer: Integrating Tree Structures into Self-Attention (D19-1)

Copied to clipboard

Challenge: Existing work on hierarchical structure in neural networks has not captured human intuitions about hierarchic structures.
Approach: They propose to add an extra constraint to attention heads of the bidirectional Transformer encoder to encourage attention heads to follow tree structures.
Outcome: The proposed model improves language modeling and learning more explainable attention scores.
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a new massive multilingual dataset is available for language modeling and machine translation training.
Approach: They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora .
Outcome: The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 .
Pretrained Models for Multilingual Federated Learning (2022.naacl-main)

Copied to clipboard

Challenge: Federated Learning (FL) is a machine learning technique that trains a model across multiple distributed clients holding local data samples, without ever storing client data in a central location.
Approach: They propose to use pretrained models to study three multilingual language tasks . they also examine impact of non-IID text on FL in naturally occurring data .
Outcome: The proposed methods perform better than centralized learning even when using non-IID partitioning.
Fusing Context Into Knowledge Graph for Commonsense Question Answering (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to combine language modeling and knowledge graphs (KG) lack the context to provide a more precise understanding of the concepts.
Approach: They propose to use external entity descriptions to provide contextual information for commonsense question answering models.
Outcome: The proposed model achieves state-of-the-art among non-generative models in OpenBookQA and is the first of its kind in the field.
Coding Textual Inputs Boosts the Accuracy of Neural Networks (2020.emnlp-main)

Copied to clipboard

Challenge: a new approach to natural language processing uses arbitrary symbols to represent meaning . Soundex, MetaPhone, NYSIIS, logogram are used as inputs for NLP .
Approach: They propose to use arbitrary symbols to represent linguistic meaning of a word . they propose to integrate codewords with text to provide more reliable inputs .
Outcome: The proposed approach outperforms state-of-the-art models on machine translation, language modeling, and part-of speech tagging.
FQuAD: French Question Answering Dataset (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in the field of language modeling have improved state-of-the-art results on many natural language processing tasks.
Approach: They propose to use a French Question Answering Dataset to track progress of French Question answering models.
Outcome: The proposed model achieves an F1 score of 92.2 and an exact match ratio of 82.1 on the test set.
On the Generalization Ability of Retrieval-Enhanced Transformers (2023.findings-eacl)

Copied to clipboard

Challenge: Recent work on retrieval-augmented language models has shown impressive results . performance gains from retrieval to a large extent originate from overlapping tokens between the database and test data, suggesting less of non-trivial generalization than previously assumed.
Approach: They propose to off-load memory from trainable weights to a retrieval database and compare it to larger models with a larger model.
Outcome: The proposed model outperforms GPT-3 and Jurassic-1 on the Pile at 4% of the model parameters.
Towards Understanding Task-agnostic Debiasing Through the Lenses of Intrinsic Bias and Forgetfulness (2024.findings-acl)

Copied to clipboard

Challenge: Debiasing Pretrained Language Models (PLMs) are task-agnostic and can be generalizable, but its impact on language modeling ability and the risk of relearning social biases remain as the two most significant challenges.
Approach: They propose a framework which can Propagate Socially-fair Debiasing to Downstream Fine-tuning to alleviate the forgetting issue of PLMs by regularizing debiased attention heads based on the PLM’s bias levels from stages of pretraining and debiase.
Outcome: The proposed framework can Propagate Socially-fair Debiasing to Downstream Fine-tuning, indicating that the ineffectiveness of debiase can be alleviated by overcoming the forgetting issue through regularizing successfully debiased attention heads based on the PLMs’ bias levels from stages of pretraining and debiases.
Residual Learning of Neural Text Generation with n-gram Language Model (2022.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show that n-gram models can achieve satisfactory performance on a large proportion of testing cases.
Approach: They propose to learn a neural LM that fits the residual between an n-gram LM and the real-data distribution.
Outcome: The proposed model achieves additional performance gains over popular standalone models on three typical language tasks.
Drop Dropout on Single Epoch Language Model Pretraining (2025.findings-acl)

Copied to clipboard

Challenge: Initial dropout was seen as a breakthrough regularization technique that reduced overfitting, yet single-epoch pretraining tasks common to modern LLMs yield minimal overfit.
Approach: They propose to use dropout during single-epoch pretraining to reduce overfitting in language modeling, morpho-syntax, question answering, and MNLI to improve performance.
Outcome: The results show that dropout is not used in large LLMs and improves performance in language modeling, morpho-syntax, question answering, and MNLI.
Unsupervised Recurrent Neural Network Grammars (N19-1)

Copied to clipboard

Challenge: RNNGs model syntax and structure by incrementally generating a syntax tree and sentence in a top-down, left-to-right order.
Approach: They explore unsupervised learning of recurrent neural network grammars for language modeling and grammar induction.
Outcome: The proposed model outperforms standard sequential language models and improves parsing performance.
Implicit n-grams Induced by Recurrence (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies show that self-attention based models have limitations on modeling sequential transformations.
Approach: They propose to extract some explainable features from trained RNNs that are reminiscent of classical n-grams features.
Outcome: The proposed models can model interesting linguistic phenomena such as negation and intensification.
Hierarchical Transformers Are More Efficient Language Models (2022.findings-naacl)

Copied to clipboard

Challenge: Transformers are impressive but inefficient and costly, which limits their applications and accessibility.
Approach: They first use different ways to downsample and upsamplify activations in Transformers to make them hierarchical.
Outcome: The proposed model outperforms Transformers on the ImageNet32 and enwik8 benchmarks.
Knowledge-Augmented Language Model and Its Application to Unsupervised Named-Entity Recognition (N19-1)

Copied to clipboard

Challenge: Current language models are unable to efficiently model entity names observed in text providing insufficient context.
Approach: They propose to augment a traditional model with an external knowledge base to model entity names observed in text.
Outcome: The proposed model improves on a Named Entity Recognition (NER) task by requiring no additional information such as named entity tags.
AwarenessBench: Assessing Cognitive Capabilities of Language Models (2026.acl-long)

Copied to clipboard

Challenge: Language models exhibit increasingly consciousness-like behaviors, requiring a baseline to evaluate their cognitive abilities.
Approach: They propose a benchmark to assess the cognitive abilities of language models (LMs) they compare 18 state-of-the-art LMs to human models in metacognition, self-awareness, social awareness and situational awareness .
Outcome: Evaluating 18 state-of-the-art LMs, they find they consistently surpass baselines . but most models fall short in metacognition and self-awareness, the study finds .
Non-Exchangeable Conformal Language Generation with Nearest Neighbors (2024.findings-eacl)

Copied to clipboard

Challenge: Existing methods to evaluate reliability of generated text are lacking in natural language generation.
Approach: They propose a non-exchangeable conformal prediction method that provides bounds on coverage . they validated their method with k-NN retrieval and show that it produces encouraging results .
Outcome: The proposed method produces encouraging results in machine translation and language modeling tasks.
Explicitly Modeling Syntax in Language Models with Incremental Parsing and a Dynamic Oracle (2021.naacl-main)

Copied to clipboard

Challenge: Failing to capture the structure of input language could lead to generalization problems and over-parametrization.
Approach: They propose a new syntax-aware language model that explicitly models the structure with an incremental parser and maintains the conditional probability setting of a standard language model.
Outcome: The proposed model can achieve strong results in language modeling, parsing, and syntactic generalization tests while using fewer parameters than other models.
An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models (2022.acl-long)

Copied to clipboard

Challenge: Recent work has shown pre-trained language models capture social biases from the large amounts of text they are trained on.
Approach: They propose to use Counterfactual Data Augmentation, Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebia as bias mitigation techniques to quantify their effectiveness.
Outcome: The proposed techniques are Counterfactual Data Augmentation (CDA), Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebia.
Contextual String Embeddings for Sequence Labeling (C18-1)

Copied to clipboard

Challenge: Recent advances in language modeling have made it viable to model language as distributions over characters.
Approach: They propose to leverage internal states of a trained character language model to produce a new type of word embeddings.
Outcome: The proposed embeddings outperform the state-of-the-art on four classic sequence labeling tasks.
Exploring Multitask Learning for Low-Resource Abstractive Summarization (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent work shows that training text encoders using data from multiple tasks helps to produce an encoder that can be used in numerous downstream tasks with minimal fine-tuning.
Approach: They incorporate four different tasks to improve abstractive summarization performance . they use a pretrained BERT model and train all tasks using a small-scale training corpus .
Outcome: The proposed model outperforms a model trained in a multitask setting with no additional summarization data.
Improved Language Modeling by Decoding the Past (P19-1)

Copied to clipboard

Challenge: Existing methods to improve language modeling performance are based on regularized LSTMs with a large number of parameters and training time.
Approach: They propose a method that decodes the last token in context using the predicted distribution of the next token.
Outcome: The proposed method improves perplexity on the Penn Treebank dataset by 1.8 points and 2.3 points on the WikiText-2 datasets.
Long-Context Language Modeling with Parallel Context Encoding (2024.acl-long)

Copied to clipboard

Challenge: Existing long-context models degenerate with retrieved contexts.
Approach: They propose a framework that can be applied to existing decoder-only LLMs for context expansion.
Outcome: The proposed framework can be applied to any existing decoder-only LLMs for context expansion.
Improving Neural Language Models by Segmenting, Attending, and Predicting the Future (P19-1)

Copied to clipboard

Challenge: Common language models typically predict the next word given a past context.
Approach: They propose a method that aligns the given context and the following phrase . they define syntactic heights and phrase segmentation rules to enable it to learn .
Outcome: The proposed model outperforms strong baseline models on Wikitext-103 dataset.
Trained on 100 million words and still in shape: BERT meets British National Corpus (2023.findings-eacl)

Copied to clipboard

Challenge: masked language models are trained on ever larger corpora, but pre-training on a modestly-sized but representative, well-balanced, and publicly available corpus can reach better performance than the original BERT model.
Approach: They propose an optimized LM architecture called LTG-BERT that can be used to train a competitive language model on a small and standardizable corpus.
Outcome: The proposed architecture outperforms the original English BERT model on a representative, well-balanced and publicly available corpus.
Rational Recurrences (D18-1)

Copied to clipboard

Challenge: Recent studies show that neural models lack strong intuitions . recent studies show connections between convolutional neural networks and weighted finite state automata (WFSAs)
Approach: They show that some recurrent neural networks share a connection to weighted finite state automata (WFSAs) they define rational recurrences as recursive hidden state update functions . they propose to use these functions to write forward calculations of a finite set of WFSA's .
Outcome: The proposed model outperforms two baselines on language modeling and text classification.
AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: Pretrained language models often need to specialize to specific domains.
Approach: They propose an approach that performs weight-space averaging of adapters trained on different domains.
Outcome: The proposed approach improves performance to new domains without extra training.
Better Character Language Modeling through Morphology (P19-1)

Copied to clipboard

Challenge: Inflected words benefit more from explicitly modeling morphology than uninflectes . morphological supervision is also used to augment character language models in low-resource languages .
Approach: They add morphological supervision to character language models via multitasking to improve BPC performance across 24 languages even when morphology data and language modeling data are disjointed.
Outcome: The addition improves performance even when morphology data and language modeling data are disjointed.
Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection (2022.emnlp-main)

Copied to clipboard

Challenge: Language models rely on massive web crawls for diverse text data, but are rife with undesirable content.
Approach: They analyze newspaper articles written by students from across the country to determine whose language is preferred by a quality filter.
Outcome: The results show that newspapers from wealthier, educated, and urban zones are more likely to be classified as high quality.
UNIFIEDQA: Crossing Format Boundaries with a Single QA System (2020.findings-emnlp)

Copied to clipboard

Challenge: Question answering (QA) tasks have been posed using a variety of formats . a new study aims to develop specialized QA models that can be used to train QA systems .
Approach: They build a pre-trained question answering model that performs well across 19 QA datasets . they argue that format-specialized models can limit the ability to teach reasoning .
Outcome: a new model that trains on QA datasets performs on par with 8 models trained on individual datasets . a single model that trained on UNIFIEDQA performs well on 19 QA data .
On Faithfulness and Factuality in Abstractive Summarization (2020.acl-main)

Copied to clipboard

Challenge: Existing conditional text generation models produce unfaithful and unfaithed summaries . current models accomplish a high level of fluency and coherence .
Approach: They propose to use pretrained models for document summarization to better understand hallucinations . they find that textual entailment measures better correlate with faithfulness .
Outcome: The proposed models generate faithful and factual summaries as evaluated by humans.
Autoregressive Pre-Training on Pixels and Texts (2024.emnlp-main)

Copied to clipboard

Challenge: pixel-based language modeling integrates visual and textual data to improve performance of language models.
Approach: They propose a method that integrates visual and textual data into an autoregressive framework.
Outcome: The proposed method improves performance of pixel-based language models by incorporating visual and textual data.
Neural Multitask Learning for Simile Recognition (D18-1)

Copied to clipboard

Challenge: Simile is a special type of metaphor, where comparators such as like and as are used to compare two objects.
Approach: They propose a neural network framework for simile sentence classification, simile component extraction and language modeling.
Outcome: The proposed framework outperforms rule-based and feature-based approaches in simile sentence classification and simile component extraction tasks.
Continual and Multi-Task Architecture Search (P19-1)

Copied to clipboard

Challenge: Recent studies have shown that architecture search can improve performance on language modeling and image classification tasks with reasonable training speed.
Approach: They propose a continual architecture search approach that continually evolves the model parameters during sequential training of several tasks without losing performance on previously learned tasks.
Outcome: The proposed approach improves language modeling and image classification with reasonable training speed and a weight-sharing strategy.
MirasText: An Automatically Generated Text Corpus for Persian (L18-1)

Copied to clipboard

Challenge: Natural language processing is one of the most important fields of artificial intelligence.
Approach: They propose to use MirasText to generate Persian text corpus from Persian websites . MiraSText has over 2.8 million documents and over 1.4 billion tokens .
Outcome: The generated corpus has over 2.8 million documents and over 1.4 billion tokens . MirasText has over 800 billion token tokens and more than 300 thousand articles .
BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing long context models suffer from performance decline when the input text exceeds their length limit.
Approach: They propose a multi-task long context benchmark to evaluate LLMs' long context ability using 10 datasets from 5 different NLP tasks.
Outcome: The proposed model covers 5 domains and core capacities of large language models.
xGQA: Cross-Lingual Visual Question Answering (2022.findings-acl)

Copied to clipboard

Challenge: a lack of multilingual multimodal datasets has hindered multimodal vision and language modeling efforts.
Approach: They propose a multilingual evaluation benchmark for the visual question answering task . they extend the established English GQA dataset to 7 typologically diverse languages .
Outcome: The proposed methods outperform current state-of-the-art models in zero-shot cross-lingual settings, but the accuracy remains low across languages.
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling (2025.acl-long)

Copied to clipboard

Challenge: Speculative sampling is an efficient way to accelerate the auto-regressive generation process of large language models.
Approach: They propose a frequency-ranked speculative sampling framework that optimizes draft candidate selection through vocabulary space compression.
Outcome: Experiments show that FR-Spec reduces LM Head computation overhead by 75% while ensuring the equivalence of the final output distribution.
Selective Differential Privacy for Language Modeling (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods to protect sensitive data from leaking are over-pessimistic and undifferentiated.
Approach: They propose a new privacy notion, selective differential privacy, to provide rigorous privacy guarantees on the sensitive portion of the data to improve model utility.
Outcome: The proposed privacy-preserving mechanism achieves better utility while remaining safe under various privacy attacks compared to baselines.
Investigating the effect of auxiliary objectives for the automated grading of learner English speech transcriptions (2020.acl-main)

Copied to clipboard

Challenge: a growing demand for the ability to communicate in English means automated tutoring and assessment systems are becoming more popular.
Approach: They propose to use automatic speech recognition transcripts to grade spontaneous speech based on textual features.
Outcome: The proposed system improves on a transformer encoder with native language identification as an auxiliary task.
MANTa: Efficient Gradient-Based Tokenization for End-to-End Robust Language Modeling (2022.findings-emnlp)

Copied to clipboard

Challenge: Subword tokenization algorithms have been an essential component of language modeling but their static nature results in important flaws that degrade the models’ downstream performance and robustness.
Approach: They propose a module for Adaptive Neural TokenizAtion that is differentiable and trained end-to-end with the language model.
Outcome: The proposed tokenizer improves robustness to character perturbations and out-of-domain data.
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal (2025.acl-long)

Copied to clipboard

Challenge: Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens.
Approach: They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions.
Outcome: The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000.
Probabilistic Robustness for Data Filtering (2023.eacl-main)

Copied to clipboard

Challenge: Modern machine learning works with massive amounts of data on a range of tasks like language modeling, object detection, and data mining.
Approach: They propose a probabilistic robustness rewarded data optimization approach to enhance the model's generalization power by selecting training data that optimizes probabilistic metrics.
Outcome: The proposed approach achieves +17.2% increase of accuracy and -28.05 decrease of perplexity on unknown-domain test sets.
Interpretable and Low-Resource Entity Matching via Decoupling Feature Learning from Decision Making (2021.acl-long)

Copied to clipboard

Challenge: Entity Matching (EM) aims at recognizing entity records that denote the same real-world object.
Approach: They propose a novel EM framework that consists of Heterogeneous Information Fusion and Key Attribute Tree Induction to decouple feature representation from matching decision.
Outcome: The proposed framework outperforms SOTA EM models on 6 public datasets and 3 industrial datasets.
Adapting Open Domain Fact Extraction and Verification to COVID-FACT through In-Domain Language Modeling (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to verify scientifically false online information are limited by the lack of training data in the scientific domain.
Approach: They propose an in-domain language modeling method for fact extraction and verification systems . they use SCIFACT to extract scientifically false online information .
Outcome: The proposed method improves accuracy 30% on SCIFACT dataset . state-of-the-art model achieves only 46.6% precision, which is hard to be trusted for users.
Learning to Generate Word Representations using Subword Information (C18-1)

Copied to clipboard

Challenge: Existing word-based approaches to learning word representations are blind to subword information in words.
Approach: They propose a character-based word representation approach to learn word representations from characters.
Outcome: The proposed model outperforms baseline models that regard words as atomic units . the proposed model achieves 18.5% improvement on average in perplexity for morphologically rich languages .
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for LRM unlearning overlook critical information leakage in reasoning traces, even when final answers are successfully removed.
Approach: They propose a method that suppresses reasoning traces while preserving the model's general reasoning ability.
Outcome: The proposed method significantly reduces reasoning trace leakage and achieves strong performance across reasoning and safety benchmarks, including WMDP, StrongReject, JBB-Behaviors and WildJailbreak.
Universal Adversarial Triggers for Attacking and Analyzing NLP (D19-1)

Copied to clipboard

Challenge: Using adversarial triggers, a model can produce a specific prediction . adversarial attacks are useful for evaluation and interpretation .
Approach: They propose a gradient-guided search over tokens that finds short adversarial triggers that successfully trigger the target prediction.
Outcome: The proposed algorithm finds short trigger sequences that successfully trigger the target prediction.
EnCBP: A New Benchmark Dataset for Finer-Grained Cultural Background Prediction in English (2022.findings-acl)

Copied to clipboard

Challenge: Existing research on cultural background modeling is coarse-grained and does not examine cultural differences among speakers of the same language.
Approach: They use a news-based cultural background prediction dataset to annotate, validate and benchmark NLP models with cultural background features.
Outcome: The proposed model improves on nine syntactic, semantic, and psycholinguistic tasks while introducing cultural background information does not improve the Go-Emotions task due to text domain conflicts.
Enabling Language Models to Fill in the Blanks (2020.acl-main)

Copied to clipboard

Challenge: Infilling is the task of predicting missing spans of text at any position in a document.
Approach: They propose a framework which can be used to infill entire sentences . they train off-the-shelf LMs on sequences containing concatenation of masked text .
Outcome: The proposed approach can infill entire sentences on short stories, scientific abstracts, and lyrics.
Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data (2022.acl-long)

Copied to clipboard

Challenge: Experimental results show that REtrieving from the traINing datA only can lead to significant gains on multiple NLG and NLU tasks.
Approach: They propose to retrieve training instances from traINing datA and concatenate them with input to generate output.
Outcome: The proposed method achieves state-of-the-art results on XSum, BigPatent, and CommonsenseQA.
ERNIE-Doc: A Retrospective Long-Document Modeling Transformer (2021.acl-long)

Copied to clipboard

Challenge: Existing models for document-level language pretraining are not suitable for long documents due to their quadratically increasing memory and time consumption.
Approach: They propose a document-level language pretraining model based on Recurrence Transformers.
Outcome: The proposed model outperforms existing models on language understanding tasks.
Bringing Emerging Architectures to Sequence Labeling in NLP (2026.eacl-long)

Copied to clipboard

Challenge: Pretrained Transformer encoders are the dominant approach to sequence labeling . however, few have been applied to sequence labels on flat or simplified tasks .
Approach: They propose to use pretrained Transformer encoders to model relations across words . they find that the architectures adapt well across tagging tasks that vary in complexity .
Outcome: The proposed architectures perform well across tagging tasks across languages and datasets.
Adapting Language Models to Compress Contexts (2023.emnlp-main)

Copied to clipboard

Challenge: Transformer-based language models have a finite context window and expensive computational cost of processing long text documents.
Approach: They propose to adapt pre-trained LMs into AutoCompressors to compress text into summary vectors . authors propose to use summary vector to speed up inference over long contexts based on a finite context window .
Outcome: The proposed model can compress long contexts into summary vectors, which are accessible as soft prompts.
Taking a Deep Breath: Enhancing Language Modeling of Large Language Models with Sentinel Tokens (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have explored compression and accumulation methods to compress contexts, but these methods lose useful context information during the compression process, leading to performance degradation.
Approach: They propose a method that allows LLMs to take a deep breath and insert a special token at the end of each chunk.
Outcome: Experiments on language modeling and out-of-domain tasks validate the superiority of the proposed method.
A Batch Normalized Inference Network Keeps the KL Vanishing Away (2020.acl-main)

Copied to clipboard

Challenge: Variational Autoencoder (VAE) is widely used to approximate a model’s posterior on latent variables.
Approach: They propose to let the Kullback–Leibler divergence individual follow a distribution across the whole dataset and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL’s distribution positive.
Outcome: The proposed approach can avoid posterior collapse effectively and efficiently without introducing any new model component or modifying the objective.
When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Multilingual language models are widely used to extend NLP systems to low-resource languages.
Approach: They pre-train over 10,000 monolingual and multilingual language models for over 250 languages including multiple language families that are under-studied in NLP.
Outcome: The results show that adding multilingual data improves low-resource language modeling performance, similar to increasing low-source dataset sizes by up to 33%.
Sentence-level Privacy for Document Embeddings (2022.acl-long)

Copied to clipboard

Challenge: Existing studies on the definitions of good privacy for natural language use argue that different applications and models require different definitions.
Approach: They propose a technique that uses a combination of statistics and language modeling to produce high (768) dimensional, general -SentDP document embeddings that guarantee a single sentence can be substituted with any other sentence.
Outcome: The proposed method outperforms baseline methods with weaker guarantees like word-level Metric DP and outperformed baseline methods.
Improving Word Embedding Factorization for Compression Using Distilled Nonlinear Neural Decomposition (2020.findings-emnlp)

Copied to clipboard

Challenge: Word-embeddings are vital components of natural language processing (NLP) but they consume a lot of memory which poses a challenge for edge deployment.
Approach: They propose an embedding compression method based on matrix decomposition and knowledge distillation that initializes weights of pre-trained word-embeddings and fine-tunes end-to-end.
Outcome: The proposed method has higher BLEU score on translation and lower perplexity on language modeling compared to complex, difficult to tune methods.
FENAS: Flexible and Expressive Neural Architecture Search (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent approaches to architecture search have shown good improvements in terms of performance with reasonable training speed.
Approach: They propose an algorithm with more activation functions, input edges, and atomic operations to search for architectures that are optimal for given task.
Outcome: The proposed algorithm reproduces well-known LSTM and GRU architectures and initializes with them for finding architectures more efficiently.
Effective Long-Context Scaling of Foundation Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are rapidly deployed and continue to evolve through scaling.
Approach: They propose a method to train strong long-context LLMs that are capable of utilizing massive context windows of up to 32,000 tokens.
Outcome: The proposed model can surpass gpt-3.5-turbo-16k's overall performance on long-context benchmarks with a cost-effective instruction tuning procedure that is free of expensive annotations.
ConVEx: Data-Efficient and Few-Shot Slot Labeling (2021.naacl-main)

Copied to clipboard

Challenge: ConVEx is an efficient pretraining and fine-tuning neural approach for slot-labeling dialog tasks.
Approach: They propose an efficient pretraining and fine-tuning neural approach for slot-labeling dialog tasks that uses a pairwise cloze task and reddit data.
Outcome: The proposed approach is well aligned with its intended use on slot-labeling tasks and can be used across a range of domains and data sets.
Is Supervised Syntactic Parsing Beneficial for Language Understanding Tasks? An Empirical Investigation (2021.eacl-main)

Copied to clipboard

Challenge: Traditional NLP has long held (supervised) syntactic parsing necessary for successful higher-level semantic language understanding (LU).
Approach: They empirically examine the usefulness of supervised parsing for semantic LU in LM-pretrained transformer networks.
Outcome: The proposed model is based on LM-pretrained transformer networks with a biaffine parsing head and fine-tuned for LU tasks.
Contextual morphologically-guided tokenization for Latin encoder models (2026.eacl-long)

Copied to clipboard

Challenge: Existing tokenization methods focus on information-theoretical goals like high compression and low fertility rather than linguistic goals like morphological alignment.
Approach: They propose to incorporate morphological knowledge into tokenization to improve both morphology and downstream performance.
Outcome: The proposed tokenization improves overall performance on four downstream tasks.
Revisiting the Hierarchical Multiscale LSTM (C18-1)

Copied to clipboard

Challenge: Hierarchical Multiscale LSTM model learns structure from character input . high complexity of architecture, training and implementations might hinder its applicability .
Approach: They propose to reproduce and ablate hierarchical multiscale LSTM language model and show that simplifying certain aspects of the architecture can improve its performance.
Outcome: The proposed model performs better when simplified and linguistic units are learned by different levels of the model.
High-Dimension Human Value Representation in Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to align large language models with human values and preferences are not able to be applied to all tasks and fields.
Approach: They propose a high-dimensional representation of symbolic human value distributions in LLMs that is orthogonal to model architecture and training data.
Outcome: The proposed representations are evaluated on 15 open-source and commercial LLMs and are self-supervised from the value-relevant output of 8 LLM models.
TriEmbed: Bridge the Gap between Text and Token Indices with Embedding Reparameterization (2025.findings-acl)

Copied to clipboard

Challenge: a current paradigm of language modeling discards linguistic relations between tokens during tokenization, creating a fundamental gap . empirical results show that TriEmbed provides more linguistically informative token embeddings .
Approach: They propose a reparameterization method that incorporates morphological relationships . they propose to organize the vocabulary into a Trie structure to reparametrize embeddings .
Outcome: Empirical results show that TriEmbed outperforms existing token embeddings while offering more linguistically informative token embeds.
Decoding Reading Goals from Eye Movements (2025.acl-long)

Copied to clipboard

Challenge: a study examines whether readers can distinguish between two types of reading goals: information seeking and ordinary reading for comprehension.
Approach: They propose a method to distinguish between two types of reading goals: information seeking and ordinary reading for comprehension.
Outcome: The proposed model solves the reading goal-oriented task with the most accurate predictions in real time, the authors say .
Chunk-based Nearest Neighbor Machine Translation (2022.emnlp-main)

Copied to clipboard

Challenge: Semi-parametric models augment generation with retrieval, but require expensive retrieval operation for every generated token.
Approach: They propose a semi-parametric model which augments generation with retrieval by retrieving tokens from a datastore.
Outcome: The proposed model can retrieve chunks of tokens from the datastore, instead of a single token, with a low decoding speed.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context (P19-1)

Copied to clipboard

Challenge: Term memory networks (RNNs) are difficult to optimize due to gradient vanishing and explosion.
Approach: They propose a neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence.
Outcome: The proposed method improves state-of-the-art performance on short and long sequences and generates coherent, novel text articles with thousands of tokens.
A Cheaper and Better Diffusion Language Model with Soft-Masked Noise (2023.emnlp-main)

Copied to clipboard

Challenge: Existing diffusion models have limitations in modeling discrete data, e.g., languages . we present a novel diffusion model for language modeling inspired by linguistic features in languages based on iterative denoising .
Approach: They propose a method that iteratively denoises and adds corruptions to the textual data through soft-masking to better noise it.
Outcome: The proposed model achieves better generation quality and lower training cost than current models with better performance.
Is Word Segmentation Necessary for Deep Learning of Chinese Representations? (P19-1)

Copied to clipboard

Challenge: Using word-based models, we compare word-oriented models with char-based ones . word-driven models are more vulnerable to data sparsity and the presence of out-of-vocabulary words .
Approach: They benchmark word-based models with char-based model which does not involve word segmentation in four NLP benchmark tasks.
Outcome: The proposed model outperforms char-based models in four NLP benchmark tasks.
Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing Diffusion Language Models rely on hard binary masking and discrete token assignments, which hinder the revision of early decisions.
Approach: They propose a diffusion-based language modeling approach that replaces hard binary masks with evolving soft token distributions.
Outcome: The proposed approach outperforms existing DLMs on multiple benchmarks.
Causal Distillation for Language Models (2022.naacl-main)

Copied to clipboard

Challenge: Distillation efforts have led to language models that are more compact and efficient without serious drops in performance.
Approach: They propose to augment distillation with a third objective that encourages the student model to imitate the causal dynamics of the teacher through a distillation interchange intervention training objective (DIITO).
Outcome: The proposed method lowers perplexity on the WikiText-103M corpus and improves on the GLUE benchmark, SQuAD, and CoNLL-2003.
Long-Range Language Modeling with Selective Cache (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models that use transformers to model language cost quadratically increase with sequence length.
Approach: They propose a selective cache which stores key-value pairs from previous contexts.
Outcome: The proposed selective cache outperforms XL cache and compressive cache by considerable margins.
Revisiting the Task of Scoring Open IE Relations (L18-1)

Copied to clipboard

Challenge: Recent Open Information Extraction systems allow us to extract ever larger (yet incomplete) open-domain Knowledge Bases from text.
Approach: They propose a baseline model which gives competitive results in a previously defined protocol and provides an independent source of signal to judge arbitrary fact plausibility.
Outcome: The proposed model gives competitive results in the previously defined protocol and provides an independent source of signal to judge arbitrary fact plausibility.
Increasing Learning Efficiency of Self-Attention Networks through Direct Position Interactions, Learnable Temperature, and Convoluted Attention (2020.coling-main)

Copied to clipboard

Challenge: SANs are an integral part of successful neural networks such as Transformer . training SAN on a task or pretraining them on language modeling requires large amounts of data and compute resources.
Approach: They propose to modify SANs to enable faster learning, i.e., higher accuracies after fewer update steps.
Outcome: The proposed modifications enable faster learning, i.e., higher accuracies after fewer update steps.
QUDeval: The Evaluation of Questions Under Discussion Discourse Parsing (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics poorly approximate parser quality, says a new study . questions under discussion is a linguistic framework that views discourse as asking questions and answering them .
Approach: They propose a framework for automatic evaluation of QUD parsing . they use a dataset of fine-grained evaluation of 2,190 QUD questions .
Outcome: The proposed framework shows that satisfying constraints of QUD is still challenging for modern LLMs.
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling (2022.naacl-main)

Copied to clipboard

Challenge: Existing models of language understanding are based on explicit representations of hierarchical structure, but there are good reasons to doubt that they can be said to understand language in any meaningful way.
Approach: They examine whether syntactic and semantic graph representations can complement and improve neural language modeling.
Outcome: The proposed model outperforms pretrained models on English WSJ in perplexity and other metrics.
Decoding Speculative Decoding (2025.naacl-long)

Copied to clipboard

Challenge: Speculative decoding is a widely used technique to speed up inference for Large Language Models (LLMs) Autoregressive decoding has been known to be hardware inefficient, leading to poor resource utilization and low throughput during inference.
Approach: They propose to use a draft model to generate speculative tokens and then use the target LLM to verify those tokens.
Outcome: The proposed model can provide 111% higher throughput than existing draft models and generalizes further to all LLaMA models and supervised fine-tuned models.
Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German (2020.lrec-1)

Copied to clipboard

Challenge: Despite the small pool of speakers, there are still few natural language processing corpora, studies or tools for Swiss German.
Approach: They propose to use a web scraper to generate the largest Swiss German text corpus . they show that the tool can be applied to other low-resource languages as well .
Outcome: The proposed tool significantly improves language modeling in Swiss German, the authors show .
Compositional Demographic Word Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are usually derived from corpora containing text from many individuals . however, they cannot account for user-specific word preferences, such as using the same word in different ways across contexts.
Approach: They propose a new form of personalized word embeddings that use demographic-specific word representations derived compositionally from full or partial demographic information for a user.
Outcome: The proposed representations outperform generic representations on two English language tasks.
Multi-Task Learning with Language Modeling for Question Generation (D19-1)

Copied to clipboard

Challenge: Existing work on answer-aware questions generates a sentence and answer span as input . previous work on QG was mainly tackled by rule-based approach and neural-based one .
Approach: They propose to incorporate an auxiliary task of language modeling to help question generation in a hierarchical multi-task learning structure.
Outcome: The proposed model improves on SQuAD and MARCO datasets and human evaluation proves it.
An Imitation Learning Approach to Unsupervised Parsing (P19-1)

Copied to clipboard

Challenge: Unsupervised parsing is a form of reinforcement learning that improves syntactic structures but lacks interpretability due to its lack of ad hoc heuristics.
Approach: They propose an unsupervised approach that transfers syntactic knowledge to a Tree-LSTM model with discrete parsing actions.
Outcome: The proposed model outperforms existing models on the All Natural Language Inference dataset and achieves a new state of the art in terms of parsing F-score.
Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent improvements in NLP tasks can be attributed to the Transformer model.
Approach: They propose to use parameter-sharing methods to reduce parameter budgets in generative models by using sandwich-style parameter sharing and self-attentive embedding factorization.
Outcome: The proposed model outperforms the current RNN model even with significantly fewer parameters.
ELI5: Long Form Question Answering (P19-1)

Copied to clipboard

Challenge: Existing question answering datasets provide extractive or short answers, but less attention has been paid to open-ended questions that require explanations.
Approach: They present a large-scale corpus for long form question answering . they use a Reddit forum to provide elaborate answers to open-ended questions .
Outcome: The proposed model outperforms Seq2Seq, language modeling, and other models in human evaluations.
Rectified Sparse Attention for Efficient Long-Sequence Generation (2026.findings-acl)

Copied to clipboard

Challenge: Recent sparse decoding methods improve efficiency but suffer from KV cache misalignment, resulting in performance degradation.
Approach: They propose a method that combines block-sparse attention with periodic dense rectification to bound error accumulation and preserve alignment with the pretraining distribution.
Outcome: Experiments on math reasoning, language modeling, and retrieval tasks show that ReSA achieves near-lossless generation quality with significantly improved efficiency.
A Systematic Study of Compositional Syntactic Transformer Language Models (2025.acl-long)

Copied to clipboard

Challenge: Syntactic language models (SLMs) incorporate syntactical biases into Transformers . authors identify key aspects of design choices in existing models and novel variants based on experimental results .
Approach: They propose a framework that incorporates existing and new SLMs to enhance Transformers by incorporating syntactic biases.
Outcome: The proposed framework improves on existing models and novel variants across language modeling, syntactic generalization, summarization, and inference efficiency.
Clarifying Implicit and Underspecified Phrases in Instructional Text (2022.lrec-1)

Copied to clipboard

Challenge: Natural language consists of implicit and underspecified phrases, which can cause misunderstandings.
Approach: They propose to use wikiHow to extract human clarifications that resolve an implicit or underspecified phrase.
Outcome: The proposed model can be used to generate alternate clarifications, which may or may not be compatible with the human clarification.
Global Gallery: The Fine Art of Painting Culture Portraits through Multilingual Instruction Tuning (2024.naacl-long)

Copied to clipboard

Challenge: This study examines the ability of Large Language Models to encapsulate cultural nuances across diverse linguistic landscapes.
Approach: They examine the efficacy of language-specific instruction tuning and the impact of pretraining on dominant language data in Large Language Models.
Outcome: The findings highlight a nuanced landscape, with inconsistencies and biases, particularly in non-Western cultures.
From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that multilingual transformers are less effective in resource-lean scenarios and for distant languages.
Approach: They propose to use massively multilingual transformers to pretrain languages . they show that MMTs are less effective in resource-lean scenarios and distant languages if they are pre-trained via language modeling .
Outcome: The proposed model is less effective in resource-lean scenarios and for distant languages than cross-lingual word embeddings.
From Word to World: Can Large Language Models be Implicit Text-based World Models? (2026.acl-long)

Copied to clipboard

Challenge: Agentic learning increasingly hinges on interaction, yet real-world experience is expensive, limited, and often irreversible at inference time.
Approach: They propose a framework that reframes language modeling as next-state prediction under interaction.
Outcome: The proposed framework evaluates world models in text-based environments . it shows that sufficiently trained models capture coherent environment dynamics .
Improved Differentiable Architecture Search for Language Modeling and Named Entity Recognition (D19-1)

Copied to clipboard

Challenge: Neural architecture search (NAS) is a popular approach for finding new models and freeing researchers from the hard work of designing network architectures.
Approach: They propose differentiable neural architecture search methods for natural language processing . they remove the softmax-local constraint and apply it to named entity recognition .
Outcome: The proposed method outperforms strong baselines on the language modeling task.
Confounding Factors in Relating Model Performance to Morphology (2025.emnlp-main)

Copied to clipboard

Challenge: morphological differences between languages are unclear, but are often considered unimportant . confounding factors make it hard to compare results and draw conclusions, authors argue .
Approach: They propose to use token bigram metrics to predict difficulty of causal language modeling . they argue that confounding factors are contributing to the conflicting evidence .
Outcome: The proposed metrics better capture the relation between morphology and tokenization compared to word-based models.
Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of Tokens (2022.naacl-main)

Copied to clipboard

Challenge: Standard pre-trained language models do not see the characters that compose each token's string representation.
Approach: They probe the embedding layer of pretrained language models and show that models learn the internal character composition of whole word and subword tokens without seeing the characters coupled with the tokens.
Outcome: The embedding layers of RoBERTa and GPT2 hold enough information to accurately spell up to a third of the vocabulary and reach high character ngram overlap across all token types.
∞-former: Infinite Memory Transformer (2022.acl-long)

Copied to clipboard

Challenge: Several efficient transformers have been proposed, but they all have a finite memory capacity and are forced to drop old information.
Approach: They propose an unbounded long-term memory extension that extends the vanilla transformer by using a continuous-space attention mechanism to attend over the long-time memory.
Outcome: The proposed model can model arbitrarily long contexts while keeping the computation budget fixed.
PaLM: A Hybrid Parser and Language Model (D19-1)

Copied to clipboard

Challenge: Recent language models have shown strong data-fitting performance, but do not explicitly encode any notion of structural information.
Approach: They propose a hybrid parser and neural language model that adds an attention layer over text spans in the left context.
Outcome: The proposed model outperforms baseline models on language modeling and provides syntactically-informed representations of the context.
R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling (2021.acl-long)

Copied to clipboard

Challenge: Existing models with stacked layers do not explicitly model hierarchical structure of language understanding.
Approach: They propose a recursive Transformer model based on differentiable CKY style binary trees to emulate hierarchical composition process.
Outcome: The proposed model can predict words given their left and right abstraction nodes.
On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies on large-scale in-context language models have reported successful in-const zero- and few-shot learning ability.
Approach: They investigate the effects of the pretraining corpus on in-context learning in a Korean-centric model.
Outcome: The study shows that pretraining corpus size does not determine in-context learning ability . the findings suggest that in-constext learning is not always competitive .
HoneyBee: Progressive Instruction Finetuning of Large Language Models for Materials Science (2023.findings-emnlp)

Copied to clipboard

Challenge: LLaMa-based language model for materials science is first of its kind in the world .
Approach: They propose an instruction-based process for trustworthy data curation in materials science (MatSci-Instruct) they then apply this process to finetune a LLaMa-based language model targeted for materials science.
Outcome: The proposed model outperforms existing language models on materials science tasks and improves in successive stages of refinement.
The Role of n-gram Smoothing in the Age of Neural Networks (2024.naacl-long)

Copied to clipboard

Challenge: n-gram smoothing techniques were used to overcome overfitting problems in neural language models for decades.
Approach: They propose to convert any n-gram smoothing technique into a regularizer compatible with neural language models.
Outcome: The proposed regularizers outperform label smoothing on language modeling and machine translation.
Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks (2024.lrec-main)

Copied to clipboard

Challenge: Attention pruning techniques have been developed to identify and exploit sparseness . previous work has taken pioneering steps to discover and explain the sparsity in attention patterns .
Approach: They propose a framework that observes attention patterns in a fixed dataset and generates a global sparseness mask.
Outcome: The proposed approach saves 90% of computations and maintains quality of results.
Self-Distillation for Model Stacking Unlocks Cross-Lingual NLU in 200+ Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel on English NLU tasks, yet struggle to extend their NLU capabilities to underrepresented languages.
Approach: They integrate machine translation models (MT) directly into LLM backbones via sample-efficient self-distillation.
Outcome: The proposed model outperforms translation-test models on 127 low-resource languages.
Baked-in State Probing (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work shows language models trained on form can capture aspects of meaning without explicit state supervision.
Approach: They propose to use probing to "bake" state knowledge into language models . they propose to probe for underlying world state knowledge via text prompts .
Outcome: The proposed methods show that language models trained on form can capture the world state without state supervision.
ConfliBERT: A Pre-trained Language Model for Political Conflict and Violence (2022.naacl-main)

Copied to clipboard

Challenge: Traditionally, researchers used manual coding to track conflict processes worldwide, but the high costs and slow pace of domain experts make it difficult and costly to monitor complex and rapidly changing conflicts.
Approach: They propose a domain-specific pre-trained language model for conflict and political violence that can be used to train a language model from scratch and continue training.
Outcome: The proposed model outperforms BERT in conflict research.
Living Machines: A study of atypical animacy (2020.coling-main)

Copied to clipboard

Challenge: atypical animacy is the property of being alive, but discrepancies are not uncommon . a typical animate is represented as either animate or inanimate in a text .
Approach: They propose a method for determining whether an entity is represented as animate in a text . they use a nineteenth-century English text to analyze animacy .
Outcome: The proposed method improves on an established animacy dataset and a newly introduced resource.
The Impact of Depth on Compositional Generalization in Transformer Language Models (2024.naacl-long)

Copied to clipboard

Challenge: In this paper, we test the hypothesis that deeper transformers generalize more compositionally.
Approach: They propose to add layers to transformers to generalize more compositionally . they propose to fine-tune the models so that the total number of parameters is constant .
Outcome: The proposed model generalizes more compositionally than shallower models, but returns diminish . the proposed model can be made shallower without sacrificing performance .
Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models have demonstrated their capability with few-shot inference . however, in-domain demonstrations are not always available in real scenarios .
Approach: They propose unsupervised domain adaptation problem to adapt language models from source domain to target domain without any target labels.
Outcome: The proposed model performs better than baseline models on Sentiment Analysis and Named Entity Recognition tasks.
A Cognitive Regularizer for Language Modeling (2021.acl-long)

Copied to clipboard

Challenge: a uniform information density hypothesis is used to explain certain linguistic phenomena . a regularizer that encodes the UID hypothesis can be used for language training .
Approach: They propose to augment the canonical MLE objective with a regularizer that encodes UID . they find that regularization consistently improves perplexity in language models .
Outcome: The proposed hypothesis can be operationalized as an inductive bias for language modeling.
Noise Contrastive Estimation and Negative Sampling for Conditional Models: Consistency and Statistical Efficiency (D18-1)

Copied to clipboard

Challenge: Conditional models are frequently encountered in practice, but there has not been a rigorous theoretical analysis of NCE in this setting.
Approach: They propose to use a ranking-based and ranking-only method for conditional models to estimate parameter estimates.
Outcome: The proposed method avoids calculation of partition function or derivatives at each training step . it is closely related to negative sampling methods, now widely used in NLP .
Implicit Deep Latent Variable Models for Text Generation (D19-1)

Copied to clipboard

Challenge: Variational auto-encoders have been used for text generation but their representation power is limited due to two reasons.
Approach: They advocate sample-based representations of variational distributions for natural language . they further develop an LVM to directly match the aggregated posterior to the prior .
Outcome: The proposed model can be viewed as a natural extension of VAEs with a regularization of maximizing mutual information, mitigating the "posterior collapse" issue.
Revisiting Simple Neural Probabilistic Language Models (2021.naacl-main)

Copied to clipboard

Challenge: Recent advances in language modeling have been driven not only by advances in neural architectures, but also through hardware and optimization improvements.
Approach: They revisit the neural probabilistic language model (NPLM) of Bengio et al. (2003) which simply concatenates word embeddings within a fixed window and passes the result through a feed-forward network to predict the next word.
Outcome: The proposed model performs better on word-level language model benchmarks than a baseline Transformer with short input contexts but struggles to handle long-term dependencies.
Rethinking Complex Neural Network Architectures for Document Classification (N19-1)

Copied to clipboard

Challenge: Neural network models for many NLP tasks have grown increasingly complex in recent years . authors of recent papers question the necessity of such architectures and find them quite effective .
Approach: They propose to use regularization techniques borrowed from language modeling to improve model accuracy . they find that a simple biLSTM architecture with appropriate regularization yields competitive results .
Outcome: a simple biLSTM model outperforms the state-of-the-art on four benchmark datasets . authors say that improvements are not real, but are attributed to mundane reasons .
Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation (D19-1)

Copied to clipboard

Challenge: Existing approaches to train variational autoencoders (VAEs) have been proposed to alleviate the posterior collapse issue in NLP tasks.
Approach: They propose to introduce a mutual information term between the input and its latent variable to regularize the objective of the VAE.
Outcome: The proposed model performs better on three benchmark datasets and is comparable to state-of-the-art models.
StereoSet: Measuring stereotypical bias in pretrained language models (2021.acl-long)

Copied to clipboard

Challenge: Existing literature on stereotypical biases in language models is limited . current evaluations focus on measuring bias without considering language modeling ability .
Approach: They propose to measure stereotypical biases in four domains: gender, profession, race, and religion . they compare stereotypical and language modeling ability of popular models like BERT, GPT-2, RoBERTa and XLnet .
Outcome: The proposed model shows strong stereotypical biases in gender, profession, race, and religion domains.
EIT: Enhanced Interactive Transformer (2024.acl-long)

Copied to clipboard

Challenge: Existing multi-view learning models prioritize complementarity while ignoring consensus . EMHA allows for efficient modeling of global dependencies among tokens in parallel .
Approach: They propose an enhanced multi-head self-attention (EMHA) that prioritizes complementarity while ignoring consensus.
Outcome: The proposed method favors consensus among heads by introducing two models . it is superior on a wide range of language tasks with a modest increase in model size .
Guiding Attention for Self-Supervised Learning with Transformers (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that self-attention patterns in trained models contain a majority of non-linguistic regularities.
Approach: They propose a technique to allow efficient self-supervised learning with bi-directional Transformers by using an auxiliary loss function to guide attention heads to conform to such patterns.
Outcome: The proposed method achieves state-of-the-art in low-resource settings and is agnostic to pre-training objectives.
LaMemo: Language Modeling with Look-Ahead Memory (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to model long-term dependencies are limited to long texts with thousands of words.
Approach: They propose a look-ahead memory that augments the recurrence memory by attending to the right-side tokens and interpolating with the old memory states to maintain long-term information in the history.
Outcome: Experiments on widely used language modeling benchmarks show that LaMemo outperforms baseline models with recurrence memory.
Shortformer: Better Language Modeling using Shorter Inputs (2021.acl-long)

Copied to clipboard

Challenge: Existing methods require computationally expensive relative position embeddings.
Approach: They propose two methods that decrease input length to improve perplexity and perplexability.
Outcome: The proposed methods speed up training by a factor of 1.65 and reduce memory usage.
Language Modeling for Code-Switching: Evaluation, Integration of Monolingual Data, and Discriminative Training (D19-1)

Copied to clipboard

Challenge: Code-switching (CS) is a linguistic phenomenon defined as "the alternation of two languages within a single discourse, sentence or constituent."
Approach: They propose an ASR-motivated evaluation setup which is decoupled from an ASL system and the choice of vocabulary . they propose a discriminative training approach which works better than generative language modeling .
Outcome: The proposed evaluation setup is better than generative language modeling, the authors show . the proposed setup is decoupled from an ASR system and the choice of vocabulary .
Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic Change (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to improve neural language models perform poorly on emerging data.
Approach: They propose a lexical-level masking strategy to post-train a neural language model using static data from past years.
Outcome: The proposed method outperforms existing methods on two pre-trained language models, two classification tasks, and four benchmark datasets.
You Don’t Know My Favorite Color: Preventing Dialogue Representations from Revealing Speakers’ Private Personas (2022.naacl-main)

Copied to clipboard

Challenge: Social chatbots evolve rapidly with large pretrained language models.
Approach: They propose effective defense objectives to protect persona leakage from hidden states by a simple neural network.
Outcome: The proposed defense objectives reduce the attack accuracy from 37.6% to 0.5% while preserving language models’ powerful generation ability.
Trade-Offs Between Fairness and Privacy in Language Modeling (2023.findings-acl)

Copied to clipboard

Challenge: Existing research suggests that privacy preservation comes at the price of worsening biases in classification tasks.
Approach: They propose to incorporate privacy preservation and de-biasing techniques into training text generation models to investigate the trade-off between the two dimensions.
Outcome: The proposed model improves on bias detection, privacy attacks, language modeling, and performance on downstream tasks.
Pre-training with Meta Learning for Chinese Word Segmentation (2021.naacl-main)

Copied to clipboard

Challenge: Recent studies show that pre-trained models are beneficial to Chinese Word Segmentation (CWS). However, these models lack task-specific prior segmentation knowledge.
Approach: They propose a pre-trained Chinese word segmentation model MetaSeg which incorporates meta learning into a multi-criteria pre-training task.
Outcome: Empirical results show that MetaSeg can achieve new state-of-the-art performance on twelve widely-used CWS datasets and significantly improve model performance in low-resource settings.
Can Data Diversity Enhance Learning Generalization? (2022.coling-1)

Copied to clipboard

Challenge: a diversity advanced actor-critical reinforcement learning framework is used to improve NLP generalization and accuracy.
Approach: They introduce Diversity Advanced Actor-Critic reinforcement learning framework to improve NLP generalization and accuracy.
Outcome: The proposed framework outperforms domain adaptation and generalization baselines without using any target domain knowledge.
Language Modeling with Shared Grammar (P19-1)

Copied to clipboard

Challenge: Recent work on structure-aware models have shown promising results on language modeling, but how to incorporate structure knowledge on corpus without syntactic annotations remains an open problem.
Approach: They propose a neural variational language model which enables the sharing of grammar knowledge among different corpora.
Outcome: The proposed model converges significantly faster to lower perplexity on two popular benchmark datasets.
Toward Joint Language Modeling for Speech Units and Text (2023.findings-emnlp)

Copied to clipboard

Challenge: Speech and text are two major forms of human language and little effort has been made to model them together.
Approach: They propose to combine speech and text models to create mixed speech-text data by using different tokenizers and automatic metrics to evaluate how well the model mixes speech and texts.
Outcome: The proposed model improves over a speech-only baseline and shows zero-shot cross-modal transferability.
Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling (P19-1)

Copied to clipboard

Challenge: State-of-the-art models in natural language processing (NLP) often incorporate sentence encoder functions which generate a sequence of vectors intended to represent the in-context meaning of each word in an input text.
Approach: They conduct the first large-scale systematic study of candidate pretraining tasks, comparing 19 different tasks as alternatives and complements to language modeling.
Outcome: The proposed model can be used to train sentences on language modeling tasks.
Iterative Structured Knowledge Distillation: Optimizing Language Models Through Layer-by-Layer Distillation (2025.coling-main)

Copied to clipboard

Challenge: Structured pruning and knowledge distillation are often not efficient and require a fixed architecture, limiting flexibility.
Approach: They propose a method which integrates knowledge distillation and structured pruning by replacing transformer blocks with smaller, efficient versions during training.
Outcome: The proposed method outperforms L1 pruning and maintains four-fifths of performance on language modeling and commonsense reasoning tasks.
Why do language models perform worse for morphologically complex languages? (2025.coling-main)

Copied to clipboard

Challenge: Language models perform differently across languages, a new study suggests . morphological typology may explain some of the performance differences, authors say .
Approach: They propose to test morphological alignment of tokenizers, tokenization quality and disparities in dataset sizes and measurement to test this hypothesis.
Outcome: The proposed model shows that fusional languages perform better than fusionative languages . the authors suggest that morphological typology may explain some of the performance differences .
AdaPrompt: Adaptive Model Training for Prompt-based NLP (2022.findings-emnlp)

Copied to clipboard

Challenge: Prompt-based learning can tackle zero-shot and few-shot NLP tasks . authors propose a method that makes use of pre-trained language models .
Approach: They propose to map NLP tasks into natural language prompts, which are then filled by pre-trained language models.
Outcome: The proposed method outperforms standard prompt-based methods in few-shot settings.
Discrete Optimization for Unsupervised Sentence Summarization with Word-Level Extraction (2020.acl-main)

Copied to clipboard

Challenge: Sentence summarization systems that use latent space to reconstruct the source sentence are unwillingly exploited.
Approach: They propose a method that uses language modeling and semantic similarity metrics to find a high-scoring summary.
Outcome: The proposed method achieves state-of-the-art for unsupervised sentence summarization according to ROUGE scores.
REPLUG: Retrieval-Augmented Black-Box Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing retrieval-augmented language models require access to internal representations to enhance performance.
Approach: They introduce a retrieval-augmented language modeling framework that treats the language model as a black box and augments it with a tuneable retrieval model.
Outcome: The proposed framework improves performance on language modeling tasks by 6.3% and 5.1%.
An Exploration of Mamba for Speech Self-Supervised Models (2026.acl-long)

Copied to clipboard

Challenge: Mamba-based SSL models are promising for long-sequence modeling, speech unit extraction, and speech self-supervised learning.
Approach: They propose to use Mamba-based HuBERT models as an alternative to Transformer-based SSL architectures.
Outcome: The proposed models outperform Transformer-based models in language modeling tasks while showing superior performance on streaming ASR.
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)

Copied to clipboard

Challenge: a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say .
Approach: They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary .
Outcome: The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks.
IGA: An Intent-Guided Authoring Assistant (2021.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models have improved writing assistance functions such as autocomplete, but more complex and controllable writing assistants have yet to be explored.
Approach: They build an intent-guided authoring assistant that follows fine-grained author directives by specifying different writing intents.
Outcome: The proposed system generates output satisfying the author's intent and can be rephrased to their liking.
StableMoE: Stable Routing Strategy for Mixture of Experts (2022.acl-long)

Copied to clipboard

Challenge: Existing learning-to-route methods suffer from the routing fluctuation issue . with the model scale growing, training speed will go slower and memory requirements are heavy .
Approach: They propose a Mixture-of-Experts technique that can scale up the model size of Transformers with an affordable computational overhead.
Outcome: The proposed method outperforms existing learning-to-route methods on language modeling and multilingual machine translation.
What Kind of Language Is Hard to Language-Model? (P19-1)

Copied to clipboard

Challenge: a recent study suggests that language models perform poorly across languages.
Approach: They propose a model that fits a paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at least-pairwise parallel corpora.
Outcome: The proposed model is able to handle missing data and is aware of inter-sentence variation.
Pyramidal Recurrent Unit for Language Modeling (D18-1)

Copied to clipboard

Challenge: Long short term memory units are powerful tools for language modeling, but their performance can be limited by the number of parameters.
Approach: They propose a pyramidal recurrent unit which enables learning representations in high dimensional space with more generalization power and fewer parameters.
Outcome: The proposed model outperforms existing models with different gating mechanisms and transformations on word-level language modeling tasks.
Language Modeling with Sparse Product of Sememe Experts (D18-1)

Copied to clipboard

Challenge: Existing language modeling methods rely on large-scale text data to learn the sequential patterns of words.
Approach: They propose to use sememes to represent the implicit semantics behind words for language modeling . they propose to employ sememe-driven language models to fine-grained semem-level semantics .
Outcome: Experiments on language modeling and the downstream application of headline generation show the effectiveness of SDLM.
Open ASR for Icelandic: Resources and a Baseline System (L18-1)

Copied to clipboard

Challenge: Existing language resources are not sufficient for less-resourced languages, but a system with sufficient resources is needed.
Approach: They describe available language resources and their preparation for use in a large vocabulary speech recognition system for Icelandic.
Outcome: The proposed system improves on acoustic training sets and a speech corpus with a pronunciation dictionary.
The Importance of Being Recurrent for Modeling Hierarchical Structure (D18-1)

Copied to clipboard

Challenge: Recent work shows that recurrent neural networks can implicitly capture hierarchical information when trained to solve common natural language processing tasks.
Approach: They propose a convolutional sequence-to-sequence model that exploits hierarchical information implicitly.
Outcome: The proposed model is recurrent and non-recurrent, and it can model hierarchical structure implicitly.
Simple Unsupervised Summarization by Contextual Matching (P19-1)

Copied to clipboard

Challenge: Existing methods for sentence summarization require a large amount of parallel data for supervision to work.
Approach: They propose an unsupervised method for sentence summarization using only language modeling.
Outcome: The proposed method maintains continuous contextual matching while maintaining output fluency without any paired examples.
ABC: Attention with Bounded-memory Control (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to attention with bounded-memory control (ABC) have a quadratic complexity in sequence lengths, making it prohibitive for long sequences.
Approach: They propose a new abstraction that bounds memory size to improve efficiency . they propose bounded-memory control, which connects several efficient attention variants .
Outcome: The proposed approach outperforms existing approaches on language modeling, machine translation, and masked language model finetuning.
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models (2026.acl-long)

Copied to clipboard

Challenge: Masked diffusion language models have achieved significant progress in language modeling . however, the systematic analysis and empirical validation of their alignment on general tasks remains underexplored.
Approach: They propose a framework that analyzes the bias and variance of preference optimization loss and gradient based on Direct Preference Optimization.
Outcome: The proposed model outperforms its SFT-only predecessor on general benchmarks . it consistently outperformed other strong language models and ARMs on general tasks .
Constructing Word-Context-Coupled Space Aligned with Associative Knowledge Relations for Interpretable Language Modeling (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to train language models have limitations in interpretability . a Word-Context-Coupled Space (W2CSpace) is proposed to improve the performance of pre-trained models .
Approach: They propose a Word-Context-Coupled Space to replace pre-trained models with interpretable statistical logic.
Outcome: The proposed language model can achieve better performance and highly credible interpretability compared to state-of-the-art methods.
Dodo: Dynamic Contextual Compression for Decoder-only LMs (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to NLP are sparsifying attention patterns or approximating the attention computation with kernel methods.
Approach: They propose a method for dynamic contextual compression for decoder-only LMs.
Outcome: The proposed method reduces the cost of self-attention to a fraction of typical time and space.
A Measure-Theoretic Characterization of Tight Language Models (2023.acl-long)

Copied to clipboard

Challenge: Language modeling is a core task in natural language processing.
Approach: They propose to characterize leakage onto the set of infinite sequences by a measure-theoretic approach.
Outcome: The proposed language model families are tight, meaning they will not leak . the proposed language models are based on the 'sequence leakage' hypothesis .
Grammar Induction with Neural Language Models: An Unusual Replication (D18-1)

Copied to clipboard

Challenge: Recent work on latent tree learning attempts to develop models with parse-valued latent variables and train them on non-parsing tasks.
Approach: They propose a model with parse-valued latent variables and a strong latent tree learning result on constituency parsing.
Outcome: The proposed model outperforms all baselines and performs competitively with symbolic grammar induction systems.
Contrastive Deterministic Autoencoders For Language Modeling (2023.findings-emnlp)

Copied to clipboard

Challenge: Variational autoencoders (VAEs) are a popular family of generative models with wide applicability.
Approach: They propose to modify a deterministic model designed for images to avoid posterior collapse by controlling the entropy of the aggregate posterior to make it Gaussian.
Outcome: The proposed models outperform a broad range of VAE models on text generation and downstream tasks from representations while avoiding reparametrization steps.
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models exhibit reasonable multilingual abilities, despite predominantly English-centric pretraining.
Approach: They propose a framework that establishes multilingual alignment prior to language model pretraining and preserves this alignment using a code-switching strategy during pretraining.
Outcome: Experiments in a synthetic English to English-Clone setting show that PreAlign outperforms standard multilingual joint training in language modeling, zero-shot cross-lingual transfer, and cross-linguistic knowledge application.
Deduplicating Training Data Makes Language Models Better (2022.acl-long)

Copied to clipboard

Challenge: Existing language modeling datasets contain near-duplicate examples and long repetitive substrings.
Approach: They develop tools that allow us to deduplicate existing language modeling datasets . they found that over 1% of the unprompted output of language models is copied verbatim .
Outcome: The proposed tools reduce train-test overlap, which affects over 4% of validation sets, and improve model accuracy.
Language Modeling with a General Second-Order RNN (2020.lrec-1)

Copied to clipboard

Challenge: a number of RNNs update their state as the input sequence is processed . second-order RNN architectures show promising performance in language modeling .
Approach: They propose a second-order RNN architecture that generalizes existing ones . they use a Penn Treebank dataset to analyze how their different components affect performance .
Outcome: The proposed architecture generalizes existing RNNs on a Penn Treebank dataset . it shows that removing the first-order terms does not hinder performance .
A Mixture of h - 1 Heads is Better than h Heads (2020.acl-main)

Copied to clipboard

Challenge: Evidence has shown that multi-head attentive neural architectures are overparameterized.
Approach: They propose a multi-head attentive neural architecture that “reallocates” attention heads to different inputs.
Outcome: The proposed model outperforms baselines on machine translation and language modeling tasks.
Differentiable Window for Dynamic Local Attention (2020.acl-main)

Copied to clipboard

Challenge: Existing general purpose components for learning differentiable windows are hard to optimize.
Approach: They propose a new neural module and general purpose component for dynamic window selection that can enable more focused attentions over the input regions.
Outcome: The proposed approach improves on a myriad of NLP tasks including machine translation, sentiment analysis, subject-verb agreement and language modeling.
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families.
Approach: They present a set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition.
Outcome: The Bloom Library datasets cover 363 languages across 32 language families.
Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance Approach (2020.acl-main)

Copied to clipboard

Challenge: incorporating syntactic structure into language models has been a challenge since the 1990s.
Approach: They propose to use syntactic information to integrate syntastic structure into neural language models by providing ground truth parse trees as additional training signals.
Outcome: The proposed model achieves lower perplexity and better quality when ground truth parse trees are provided as training signals.
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali (2026.findings-acl)

Copied to clipboard

Challenge: Large-scale multitask benchmarks have driven rapid progress in language modeling, yet most emphasize low-resource languages like English.
Approach: They propose a benchmark for massive multitask language understanding in Bengali . they use a dataset that preserves mathematical content via MathML and a subset of questions most frequently missed by top systems to stress difficult cases.
Outcome: The proposed benchmark covers 24 model variants across 11 LLM families.
Layer-Condensed KV Cache for Efficient Inference of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Using a key-value cache, memory consumption is a bottleneck for high-throughput language models.
Approach: They propose a method that only computes and caches the KVs of a small number of layers, thus saving memory consumption and improving inference throughput.
Outcome: The proposed method achieves higher throughput and competitive performance than standard transformers and is orthogonal to existing transformer memory-saving techniques.
Exploring the Value of Personalized Word Embeddings (2020.coling-main)

Copied to clipboard

Challenge: a subset of words belonging to specific psycholinguistic categories vary more in their representations across users . combining generic and personalized word embeddings yields the best performance .
Approach: They propose personalized word embeddings and compare their performance to generic ones . they show that personalized word representations can be leveraged for improved performance .
Outcome: The proposed model can be used for authorship attribution.
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts (2024.findings-emnlp)

Copied to clipboard

Challenge: Reinforcement learning from human feedback (RLHF) is the primary method for aligning large language models with human preferences.
Approach: They propose to train an Absolute-Rating Multi-Objective Reward Model with multi-dimensional absolute-rating data.
Outcome: The proposed model outperforms the LLM-as-a-judge method on RewardBench . it achieves state-of-the-art performance on the benchmark .
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts (2024.findings-acl)

Copied to clipboard

Challenge: Neural architecture search (NAS) uses weight-sharing supernets to generate diverse subnetworks without retraining.
Approach: They propose a weight-sharing supernet that leverages mixture-of-experts to enhance supernet model expressiveness with minimal training overhead.
Outcome: The proposed method achieves state-of-the-art (SoTA) performance in NAS for fast machine translation models, surpassing NAS-BERT and AutoDistil across various model sizes.
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization (2024.findings-acl)

Copied to clipboard

Challenge: Recent approaches to language model alignment assume homogeneous human preferences, but actual human preferences vary widely and are hard to satisfy with a single language model.
Approach: They propose an RL-free extension of Direct Preference Optimization (DPO) that folds language modeling directly into reward modeling and trains language models as collective reward models that combine all objectives with specific weights.
Outcome: The proposed method matches or outperforms existing methods in safety alignment and long-form question answering.
On the importance of pre-training data volume for compact language models (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models.
Approach: They investigate the impact of pre-training data volume on compact language models . they use a French question answering task to train models with as little as 100 MB of text .
Outcome: The results show that pre-training data volume can improve models with as little as 100 MB of text . the results suggest that the model performance is poorer with less data than with larger datasets .
Exploring and Predicting Transferability across NLP Tasks (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in NLP demonstrate the effectiveness of training large-scale language models and transferring them to downstream tasks.
Approach: They conduct an extensive study of the transferability between 33 NLP tasks across three broad classes of problems.
Outcome: The proposed model can improve performance even with low-data source tasks that differ substantially from the target task.
Unsupervised Domain Adaptation of Language Models for Reading Comprehension (2020.lrec-1)

Copied to clipboard

Challenge: State-of-the-art reading comprehension models do not have general linguistic intelligence . accuracy of out-domain datasets is affected by the distribution of data .
Approach: They propose to use supervised RC training data in the source domain and unlabeled passages in the target domain to adapt models.
Outcome: The proposed model outperforms the model without domain adaptation with five datasets in different domains.
Resolving Indirect Referring Expressions for Entity Selection (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in language modeling have enabled new conversational systems.
Approach: They propose to use a dataset of indirect referring expressions to solve the problem of reference resolution when people use natural expressions . they propose to model the problem using 42K indirect referred expressions across three domains and a public dataset of entity pairs and utterances.
Outcome: The proposed models achieve 82%-87% accuracy in realistic settings, while reasonable invites further advances.
Prompt-based Distribution Alignment for Domain Generalization in Text Classification (2022.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models (PLMs) have achieved competitive performance on a range of NLP tasks.
Approach: They propose to learn distributional invariance across source domains via alignment regularization loss functions to improve domain generalization by prompting.
Outcome: Experiments on sentiment analysis and natural language inference show the effectiveness of the proposed method and achieve state-of-the-art results.
Structured Pruning for Efficient Generative Pre-trained Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Large-scale generative Pre-trained Language Models (PLMs) are limited in their deployment in real-world applications.
Approach: They propose to prune the feed-forward networks of generative pre-trained language models to smaller widths without designing extra operators.
Outcome: The proposed method achieves 1.51x/6.96x inference speedup on GPU/CPU with 67% size reduction.
XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection (2024.findings-acl)

Copied to clipboard

Challenge: XMoE leverages small experts and a threshold-based router to selectively engage only essential parameters.
Approach: They propose a novel MoE that leverages small experts to selectively engage only essential parameters.
Outcome: The proposed model can reduce computation load at MoE layers by over 50% without sacrificing performance.
Lifelong Event Detection via Optimal Transport (2024.emnlp-main)

Copied to clipboard

Challenge: Continual event detection (CED) is a challenging task due to catastrophic forgetting, where learning new tasks hampers performance on previous ones.
Approach: They propose a method that leverages optimal transport principles to align the optimization of a classification module with the intrinsic nature of each class, as defined by their pre-trained language modeling.
Outcome: The proposed method outperforms state-of-the-art methods on MAVEN and ACE datasets and is a pioneering solution in continual event detection.
HoLM: Analyzing the Linguistic Unexpectedness in Homeric Poetry (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on the authorship of the Homeric poems has only been done at the level of lengthier excerpts, but not individual verses, at which most suspected interpolations occur.
Approach: They present a corpus of Homeric verses with a score quantifying linguistic unexpectedness based on Perplexity.
Outcome: The proposed corpus of Homeric verses is complemented with a score quantifying linguistic unexpectedness based on Perplexity.
How Do Hyenas Deal with Human Speech? Speech Recognition and Translation with ConfHyena (2024.lrec-main)

Copied to clipboard

Challenge: Currently, attention-based models face computational hurdles in processing long sequences due to its quadratic complexity.
Approach: They propose a conformer whose encoder self-attentions are replaced with Hyena for speech processing . they propose 'confhyena' model that reduces training time by 27% at minimal cost .
Outcome: The proposed model reduces training time by 27% at the cost of minimal quality degradation.
A Comprehensive Evaluation of Quantization Strategies for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Quantization studies have focused on instruction-tuned LLMs, leaving their performance on other benchmarks unclear.
Approach: They propose a framework to evaluate quantized large language models using four dimensions . they propose to reduce the bits needed for model weights or activations with minimal performance loss .
Outcome: The proposed framework can retain comparable performance to non-quantized LLMs on most benchmarks.
Function Words as Statistical Cues for Language Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing studies have argued that function words aid learning abstract grammatical knowledge from linear input.
Approach: They examine the statistical distribution of function words and their properties . they show that function words are reliable, diverse, and informative .
Outcome: The results show that function words preserve high frequency, reliable syntactic association, phrase-boundary alignment and are informative to structural dependency.
HiRoPE: Length Extrapolation for Code Models Using Hierarchical Position (2024.acl-long)

Copied to clipboard

Challenge: Existing LLMs are constrained by their pre-trained context lengths, leading to performance issues . elucidating this limitation, we propose a training-free solution to the context length limitation in LLM applications .
Approach: They propose a method that integrates hierarchical rotary position embedding into LLMs without extra training costs.
Outcome: The proposed method improves performance on language modeling and long code completion tasks.
Prompt-Based Bias Calibration for Better Zero/Few-Shot Learning of Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Prompt-based learning is susceptible to intrinsic bias present in pre-trained language models (LMs), leading to sub-optimal performance in prompt-based zero/few-shot settings.
Approach: They propose a null-input prompting method to calibrate intrinsic bias encoded in pre-trained language models (LMs) they leverage a diverse set of auto-selected null meaning inputs generated from GPT-4 to probe intrinsic bias.
Outcome: The proposed method significantly improves zero/few-shot learning performance of LMs for both in-context learning and prompt-based fine-tuning (on average 9% and 2%, respectively).
Accelerating Toeplitz Neural Network with Constant-time Inference Complexity (2023.emnlp-main)

Copied to clipboard

Challenge: Toeplitz Neural Networks outperform commonly used Transformer-based models while benefiting from log-linear space-time complexities.
Approach: They propose to convert TNNs to SSMs during inference to combine strengths of TNN and SSM approaches.
Outcome: The proposed method outperforms most Transformer-based models while retaining the advantage of constant inference complexity.
Value-aware Approximate Attention (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approximations of dot-product attention ignore the value vectors . a value-aware objective outperforms an optimal approximate that ignores values .
Approach: They propose an approximation of a value-aware objective that substantially outperforms an optimal approximate that ignores values.
Outcome: The proposed value-aware objective outperforms an optimal approximation that ignores values in the context of language modeling.
PersonaLM: Language Model Personalization via Domain-distributed Span Aggregated K-Nearest N-gram Retrieval Augmentation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing language modeling tools for automatic speech recognition (ASR) are difficult to personalize.
Approach: They propose a domain-distributed Span-Aggregated K-nearest N-gram retrieval augmentation to improve language modeling for automatic speech recognition (ASR) personalization.
Outcome: The proposed model outperforms baselines on Wikitext-103, UserLibri, and ASAP datasets with a 10-16% improvement in perplexity and a 5-8% reduction in word error rates.
Consonant is all you need: a compact representation of English text for efficient NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: In natural language processing, the representation of text plays a crucial role in various tasks such as language modeling, sentiment analysis, and machine translation.
Approach: They propose a method to represent English text with only consonants that is more discriminative than vowels and a technique to retrieve vowel information from it.
Outcome: The proposed representation significantly reduces the overall memory and compute footprint required for storing and processing textual data.
Position Paper: MeMo: Towards Language Models with Associative Memory Mechanisms (2025.findings-acl)

Copied to clipboard

Challenge: Memorization is a fundamental ability of Transformer-based Large Language Models, achieved through learning.
Approach: They propose an architecture that explicitly memorizes sequences of tokens in layered associative memories.
Outcome: The proposed architecture shows that memorization is a fundamental ability of large language models, achieved through learning.
Exploring Quantization for Efficient Pre-Training of Transformer Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Quantization has proven to be effective after pre-training and during fine-tuning, but its effects on pre-trainer performance have remained unexplored.
Approach: They propose a linear quantization strategy to be applied during the pre-training of Transformers to improve model efficiency and stability.
Outcome: The proposed method improves model efficiency, stability, and performance while maintaining language modeling ability.
Rationales for Sequential Predictions (2021.emnlp-main)

Copied to clipboard

Challenge: Sequence models produce accurate predictions, but their decision making processes are hard to explain.
Approach: They propose an efficient algorithm to approximate sequential objective by identifying the most faithful rationales.
Outcome: The proposed algorithm is best at optimizing the sequential objective and provides the most faithful rationales.
Studying word order through iterative shuffling (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work on large language models has made this hypothesis popular . but, word order is not important enough to make sentence structure relevant .
Approach: They propose an efficient procedure that finds word order having highest likelihood under a fixed language model.
Outcome: The proposed procedure can be used to find the ordering of a bag of words having the highest likelihood under a fixed language model.
A Length-Extrapolatable Transformer (2023.acl-long)

Copied to clipboard

Challenge: Existing Transformers can only deal with the in-distribution size of inputs.
Approach: They propose a relative position embedding to explicitly maximize attention resolution . they also use blockwise causal attention during inference for better resolution a .
Outcome: The proposed model achieves strong performance in interpolation and extrapolation settings.
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models (2023.emnlp-main)

Copied to clipboard

Challenge: Multilingual models have been released, but many of the world's languages are not covered.
Approach: They propose a method that initializes the embedding matrix for a new tokenizer based on information in the source model's embeddable matrix.
Outcome: The proposed method outperforms random initialization and previous work on language modeling and on a range of downstream tasks (NLI, QA, and NER).
Types of Out-of-Distribution Texts and How to Detect Them (2021.emnlp-main)

Copied to clipboard

Challenge: Current NLP models produce unreliable or catastrophic predictions when training and test distributions differ . current models tend to produce unreliability or even catastrophic predictions that hurt user trust.
Approach: They categorize examples as exhibiting a background shift or semantic shift and use calibration and density estimation methods to detect OOD examples.
Outcome: The proposed methods beat calibration methods in background shift settings and perform worse in semantic shift settings.
Jump to Conclusions: Short-Cutting Transformers with Linear Transformations (2024.lrec-main)

Copied to clipboard

Challenge: Transformer-based language models create hidden representations of inputs at every layer, but only use final-layer representations for prediction.
Approach: They propose a method for casting hidden representations as final representations, bypassing transformer computation in-between.
Outcome: The proposed method produces more accurate predictions from hidden layers across various model scales, architectures, and data distributions.
Towards Unifying Multi-Lingual and Cross-Lingual Summarization (2023.acl-long)

Copied to clipboard

Challenge: Existing work on multilingual summarization and cross-lingual summmarization has been limited due to their different definitions.
Approach: They propose to unify MLS and CLS into a more general setting, i.e. many-to-many summarization.
Outcome: The proposed model outperforms the state-of-the-art models in the zero-shot directions.
DOS: Dependency-Oriented Sampler for Masked Diffusion Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing decoding strategies for pre-trained MDLMs rely on token-level uncertainty criteria, while largely overlooking sequence-level information and inter-token dependencies.
Approach: They propose a training-free decoding strategy that leverages inter-token dependencies to inform token updates during generation.
Outcome: Empirical results show that the proposed approach consistently achieves superior performance on both code generation and mathematical reasoning tasks.
Subspace Chronicles: How Linguistic Information Emerges, Shifts and Interacts during Language Model Training (2023.findings-emnlp)

Copied to clipboard

Challenge: Contemporary advances in NLP are built on the representational power of latent embedding spaces learned by self-supervised language models (LMs).
Approach: They use a new information theoretic probing suite to analyze representational subspaces in language models.
Outcome: The proposed approach compared performance of nine tasks across 2M pre-training steps and five seeds.
Tokenization with Factorized Subword Encoding (2023.findings-acl)

Copied to clipboard

Challenge: Subword tokenization methods are often used to project subwords onto triplets . a typical tokenizer consists of 10 000s of subword mapped onto a single index .
Approach: They propose a subword tokenization method that factorizes subwords onto triplets using a VQ-VAE model.
Outcome: The proposed tokenization method is more appropriate and robust for morphological tasks than the commonly used byte-pair encoding (BPE) tokenization algorithm.
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior work on compression prioritizes preserving perplexity, which is analogous to training loss.
Approach: They examine the impact of model compression along four dimensions: degeneration harm, representational harm, dialect bias, and language modeling and downstream task performance.
Outcome: The proposed compression methods can lead to unexpected consequences, the authors show . quantization preserves bias while pruning degrades quickly.
Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language Model (2023.emnlp-main)

Copied to clipboard

Challenge: Large and sparse feed-forward layers (S-FFN) have proven effective in scaling up the model size for pretraining large language models.
Approach: They compare S-FFN architectures for language modeling and compare their performance and efficiency . they found a simpler selection method that selects blocks through their mean aggregated hidden states .
Outcome: The proposed model size and selection method achieve lower perplexity in language model pretraining compared to existing MoE architectures.
VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show VocalNet outperforms existing open-source speech LLMs despite limited training data.
Approach: They propose a scalable and model-agnostic training framework and a novel multi-token prediction paradigm for speech generation.
Outcome: The proposed model outperforms open-source speech LLMs while outperforming existing open-sourced models.
Improving Input-label Mapping with Demonstration Replay for In-context Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: In-context learning (ICL) is an emerging capability of large autoregressive language models where a few demonstrations are appended to the input to enhance the model’s understanding of downstream NLP tasks without directly adjusting the model parameters.
Approach: They propose a method where a few demonstrations are appended to the input to enhance the model's understanding of downstream NLP tasks without directly adjusting the model parameters.
Outcome: The proposed method significantly improves the input-label mapping in ICL demonstrations.
DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing sparsification methods like pruning can lose model knowledge through parameter removal.
Approach: They propose a novel approach that achieves sparsification by partitioning pre-trained FFN layers into computational blocks.
Outcome: The proposed approach achieves superior performance across language modeling and downstream tasks under equivalent computational constraints.
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: Current research on bias in language models focuses on data quality, not temporal influences of data.
Approach: They propose a methodology to interpret the interaction between training data and model architecture in bias propagation during language modeling.
Outcome: The proposed method analyzes the interaction between training data and model architecture in bias propagation during language modeling.
s1: Simple test-time scaling (2025.emnlp-main)

Copied to clipboard

Challenge: OpenAI’s o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts.
Approach: They curate a small dataset s1K with 1,000 reasoning questions based on three criteria we validate through ablations: difficulty, diversity, and quality.
Outcome: The proposed model exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24).
Can Retriever-Augmented Language Models Reason? The Blame Game Between the Retriever and the Language Model (2023.findings-emnlp)

Copied to clipboard

Challenge: kNN-LM, REALM, DPR + FiD, Contriever + ATLAS, and Contriver + Flan-T5 are popular retriever-augmented language models for a variety of tasks.
Approach: They evaluate the strengths and weaknesses of kNN-LM, REALM, DPR + FiD, Contriever + ATLAS, and Contriver + Flan-T5 in reasoning over retrieved statements across different tasks.
Outcome: The proposed models do not exhibit strong reasoning even when provided with only the required statements.
An Empirical Analysis of the Writing Styles of Persona-Assigned LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Recent efforts to "personalize" large language models by assigning them specific personas are limited by current knowledge of how well they perform.
Approach: They use a style embedding model to analyze writing styles of persona-assigned LLMs . they find significant style differences between personas using Kullback-Leibler divergence .
Outcome: The proposed model shows significant differences in writing styles among personas across socio-demographic groups.
ProMALex: Progressive Modular Adapters for Multi-Jurisdictional Legal Language Modeling (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to training language models for each jurisdiction fail to leverage common legal principles beneficial for low-resource settings or risk negative interference from conflicting jurisdictional interpretations.
Approach: They propose a parameter-efficient framework that derives hierarchical relationships across jurisdictions and progressively inserts adapter modules across model layers based on jurisdictional similarity.
Outcome: The proposed framework outperforms fully shared and jurisdiction-specific models on two legal language modeling benchmarks.
SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation (2026.findings-acl)

Copied to clipboard

Challenge: Autoregressive (AR) language models are a dominant paradigm in the field of parallelism and non-causal modeling.
Approach: They propose a blockwise discrete diffusion model that preserves AR-compatible serving while enabling parallel intra-block generation.
Outcome: The proposed model achieves theoretical speedups over 5 and wall-clock speedup of 2.3 on H200 GPUs in latency-critical regimes.
Unsupervised Morphological Tree Tokenizer (2025.findings-acl)

Copied to clipboard

Challenge: Conventional statistical tokenizers often disrupt constituent boundaries within words, thereby corrupting semantic information.
Approach: They propose a method that uses morphological structure guidance to induce character-level structures of words by training a deep model.
Outcome: Empirical results show that the proposed method retains complete morphemes and outperforms existing methods on morphological segmentation and language modeling tasks.
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training (2025.findings-acl)

Copied to clipboard

Challenge: Existing defenses for large language models do not account for the sequential nature of text data.
Approach: They propose a lightweight yet effective empirical privacy defense that leverages token-specific characteristics to protect training data of large language models.
Outcome: The proposed approach provides strong protection against membership inference attacks and improves language modeling performance by 10% across different LLM architectures and datasets compared to baselines.
SLlama: Parameter-Efficient Language Model Architecture for Enhanced Linguistic Competence Under Strict Data Constraints (2025.emnlp-main)

Copied to clipboard

Challenge: Large-scale language models (LLMs) have shown remarkable performance across a wide array of tasks.
Approach: They propose an architecture that preserves parameter efficiency of tied models without sacrificing representational benefits of untied embeddings.
Outcome: The proposed architecture achieves a 31.72% improvement in linguistic knowledge acquisition over the baseline model.
ToMMeR - Efficient Entity Mention Detection from Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to detect text spans that refer to entities are often conflated with entity typing in a single joint task.
Approach: They propose a lightweight model that probes mention detection capabilities from early LLM layers.
Outcome: The proposed model achieves 93% recall zero-shot with 90% precision under human-calibrated LLM-judge protocol .
Sequence Reducible Holdout Loss for Language Model Pretraining (2024.lrec-main)

Copied to clipboard

Challenge: Data selection techniques have shown empirical benefits in reducing the number of gradient steps to train neural models.
Approach: They propose to modify an existing data selection technique to adapt it to the sequence losses typical in language modeling.
Outcome: The proposed technique reduces the number of steps required to train neural models by 4.3% and improves generalization ability on out of domain datasets.
Exploration-Driven Reinforcement Learning for Expert Routing Improvement in Mixture-of-Experts Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: MoE-based LLMs are not explicitly supervised to select suitable experts.
Approach: They propose Exploration-Driven Reinforcement Learning (ERL) which explicitly optimizes the router by exploration of alternative routing paths.
Outcome: The proposed method improves summarization (SAMSum, XSUM, question answering, and language modeling), and raises routing quality, delivering 8.9 higher MRR than baselines over 100 perturbed routing paths.
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing static vocabulary pruning designs that reduce memory usage suffer from rigid, one-size-fits-all designs that cause information loss during the prefill stage and lack flexibility.
Approach: They propose a decoupled dynamic vocabulary selection framework that addresses memory constraints through offloading embedding and implements a hybrid static-dynamic vocabulary selection strategy for LM Head.
Outcome: The proposed framework reduces memory usage by 99% with minimal or no degradation in performance.
LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder (2025.emnlp-main)

Copied to clipboard

Challenge: Prior research on linguistic mechanisms of large language models is limited by coarse granularity, limited analysis scale, and narrow focus.
Approach: They propose a framework for analyzing the linguistic mechanisms of large language models based on Sparse Auto-Encoders.
Outcome: The proposed framework extracts Chinese and English linguistic features across four dimensions . it uncovers intrinsic representations of linguistic knowledge in LLMs and can control outputs .
GiLT: Augmenting Transformer Language Models with Dependency Graphs (2026.acl-long)

Copied to clipboard

Challenge: Recent work focuses on syntactic tree structures of languages, in particular constituency tree structures.
Approach: They propose a Graph-Infused Layers Transformer Language Model which leverages dependency graphs to augment Transformer language models.
Outcome: The proposed model achieves better syntactic generalization while maintaining competitive perplexity compared with baseline models.
Exploring morphology-aware tokenization: A case study on Spanish language modeling (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study shows that subword tokenization improves performance of neural language models.
Approach: They propose a linguistically grounded approach to train a tokenizer on morphologically segmented data.
Outcome: The proposed tokenizer improves on a Spanish language model with morphological information.
LETS-C: Leveraging Text Embedding for Time Series Classification (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in language modeling have shown promising results when applied to time series data.
Approach: They propose a method to fine-tune large language models for time series classification tasks using text embedding models and a simple classification head.
Outcome: The proposed model outperforms the current SOTA model on a time series classification benchmark and uses only 14.5% of the trainable parameters.
LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for implementing large language models are limited by high computational and memory requirements.
Approach: They propose a lightweight binarization framework that achieves effective W(1+1)A4 quantization through a novel three-stage quantization strategy.
Outcome: The proposed framework surpasses state-of-the-art methods on W2A4 quantization settings across languages.
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Structured pruning is a practical approach to deploying large language models (LLMs) but it fails to capitalize on modest task-specific calibration signals, causing limited downstream gains.
Approach: They propose a method that removes attention heads and MLP channels using loss-based important scores . they use perplexity for language modeling and a margin-based objective for decision-style tasks .
Outcome: The proposed method lowers perplexity and improves accuracy at higher sparsity . it also stabilizes accuracy and mitigates perxity collapse without fine-tuning .
Characterizing the Expressivity of Local Attention in Transformers (2026.acl-long)

Copied to clipboard

Challenge: Existing studies show that global and local attention are expressively complementary.
Approach: They propose to restrict global attention to a fixed-size window of preceding tokens . they also propose to add local attention to local-only transformers to increase model quality .
Outcome: The proposed model outperforms the global–local transformers on formal language recognition and natural language modeling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations